• bitcoinBitcoin(BTC)$86,040.001.27%
  • ethereumEthereum(ETH)$2,743.400.18%
  • tetherTether(USDT)$1.000.00%
  • binancecoinBNB(BNB)$787.370.38%
  • rippleXRP(XRP)$1.544.08%
  • usd-coinUSDC(USDC)$1.000.01%
  • solanaSolana(SOL)$117.240.76%
  • tronTRON(TRX)$0.3489001.28%
  • zcashZcash(ZEC)$1,504.22-0.10%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.010.00%
  • HyperliquidHyperliquid(HYPE)$94.95-0.57%
  • dogecoinDogecoin(DOGE)$0.0985595.01%
  • moneroMonero(XMR)$569.870.24%
  • whitebitWhiteBIT Coin(WBT)$86.500.31%
  • chainlinkChainlink(LINK)$12.97-0.77%
  • RainRain(RAIN)$0.013548-4.38%
  • USDSUSDS(USDS)$1.00-0.01%
  • cardanoCardano(ADA)$0.2476983.08%
  • leo-tokenLEO Token(LEO)$8.980.33%
  • stellarStellar(XLM)$0.2123151.89%
  • nearNEAR Protocol(NEAR)$4.504.75%
  • uniswapUniswap(UNI)$8.83-2.55%
  • bitcoin-cashBitcoin Cash(BCH)$268.681.89%
  • Ethena USDeEthena USDe(USDE)$1.00-0.02%
  • avalanche-2Avalanche(AVAX)$10.92-2.92%
  • litecoinLitecoin(LTC)$60.441.39%
  • CantonCanton(CC)$0.1173190.41%
  • daiDai(DAI)$1.000.01%
  • USD1USD1(USD1)$1.00-0.03%
  • suiSui(SUI)$1.021.72%
  • hedera-hashgraphHedera(HBAR)$0.0948846.25%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.430.14%
  • BittensorBittensor(TAO)$317.7313.38%
  • shiba-inuShiba Inu(SHIB)$0.0000065.66%
  • crypto-com-chainCronos(CRO)$0.0664234.80%
  • Global DollarGlobal Dollar(USDG)$1.000.01%
  • MemeCoreMemeCore(M)$1.33-13.06%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.01%
  • tether-goldTether Gold(XAUT)$4,322.30-0.61%
  • okbOKB(OKB)$122.830.90%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • Ripple USDRipple USD(RLUSD)$1.00-0.01%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.14-0.31%
  • BitwayBitway(BTW)$0.82-3.57%
  • aaveAave(AAVE)$141.86-3.48%
  • mantleMantle(MNT)$0.652.75%
  • EthenaEthena(ENA)$0.211520-4.02%
  • Pump.funPump.fun(PUMP)$0.0045251.53%
  • OndoOndo(ONDO)$0.433170-1.37%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Together AI Unveils Revolutionary Inference Stack: Setting New Standards in Generative AI Performance

July 21, 2024
in AI & Technology
Reading Time: 4 mins read
A A
Together AI Unveils Revolutionary Inference Stack: Setting New Standards in Generative AI Performance
ShareShareShareShareShare

Together AI has unveiled a groundbreaking advancement in AI inference with its new inference stack. This stack, which boasts a decoding throughput four times faster than the open-source vLLM, surpasses leading commercial solutions like Amazon Bedrock, Azure AI, Fireworks, and Octo AI by 1.3x to 2.5x. The Together Inference Engine, capable of processing over 400 tokens per second on Meta Llama 3 8B, integrates the latest innovations from Together AI, including FlashAttention-3, faster GEMM and MHA kernels, and quality-preserving quantization, as well as speculative decoding techniques.

Additionally, Together AI has introduced the Together Turbo and Together Lite endpoints, starting with Meta Llama 3 and expanding to other models shortly. These endpoints offer enterprises a balance of performance, quality, and cost-efficiency. Together Turbo provides performance that closely matches full-precision FP16 models, making it the fastest engine for Nvidia GPUs and the most accurate, cost-effective solution for building generative AI at production scale. Together Lite endpoints leverage INT4 quantization for the most cost-efficient and scalable Llama 3 models available, priced at just $0.10 per million tokens, which is six times lower than GPT-4o-mini.

YOU MAY ALSO LIKE

OpenAI Faces Lawsuit From British Columbia Over Tumbler Ridge Shooting

NVIDIA Introduces SoL-Pi: Auto-Research Loops That Cut Coding Agent Token Traffic by Up to 49%

The new release includes several key components:

  • Together Turbo Endpoints: These endpoints offer fast FP8 performance while maintaining quality that closely matches FP16 models. They have outperformed other FP8 solutions on AlpacaEval 2.0 by up to 2.5 points. Together Turbo endpoints are available at $0.18 for 8B and $0.88 for 70B models, which is 17 times lower in cost than GPT-4o.
  • Together Lite Endpoints: Utilizing multiple optimizations, these endpoints provide the most cost-efficient and scalable Llama 3 models with excellent quality relative to full-precision implementations. The Llama 3 8B Lite model is priced at $0.10 per million tokens.
  • Together Reference Endpoints: These provide the fastest full-precision FP16 support for Meta Llama 3 models, achieving up to 4x faster performance than vLLM.
  • The Together Inference Engine integrates numerous technical advancements, including proprietary kernels like FlashAttention-3, custom-built speculators based on RedPajama, and the most accurate quantization techniques on the market. These innovations ensure leading performance without sacrificing quality. Together Turbo endpoints, in particular, provide up to 4.5x performance improvement over vLLM on Llama-3-8B-Instruct and Llama-3-70B-Instruct models. This performance boost is achieved through optimized engine design, proprietary kernels, and advanced model architectures like Mamba and Linear Attention techniques.
  • Cost efficiency is another major advantage of the Together Turbo endpoints, which offer more than 10x lower costs than GPT-4o and significantly reduce costs for customers hosting their dedicated endpoints on the Together Cloud. On the other hand, Together Lite endpoints provide a 12x cost reduction compared to vLLM, making them the most economical solution for large-scale production deployments.
  • The Together Inference Engine continuously incorporates cutting-edge innovations from the AI community and Together AI’s in-house research. Recent advancements like FlashAttention-3 and speculative decoding algorithms, such as Medusa and Sequoia, highlight the ongoing optimization efforts. Quality-preserving quantization ensures that even with low precision, the performance and accuracy of models are maintained. These innovations offer the flexibility to scale applications with the performance, quality, and cost-efficiency that modern businesses demand. Together AI looks forward to seeing the incredible applications that developers will build with these new tools.

Check out the Detail. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter and join our Telegram Channel and LinkedIn Group. If you like our work, you will love our newsletter..

Don’t Forget to join our 46k+ ML SubReddit

Find Upcoming AI Webinars here


Shreya Maji is a consulting intern at MarktechPost. She is pursued her B.Tech at the Indian Institute of Technology (IIT), Bhubaneswar. An AI enthusiast, she enjoys staying updated on the latest advancements. Shreya is particularly interested in the real-life applications of cutting-edge technology, especially in the field of data science.


Credit: Source link

ShareTweetSendSharePin

Related Posts

OpenAI Faces Lawsuit From British Columbia Over Tumbler Ridge Shooting
AI & Technology

OpenAI Faces Lawsuit From British Columbia Over Tumbler Ridge Shooting

September 22, 2026
NVIDIA Introduces SoL-Pi: Auto-Research Loops That Cut Coding Agent Token Traffic by Up to 49%
AI & Technology

NVIDIA Introduces SoL-Pi: Auto-Research Loops That Cut Coding Agent Token Traffic by Up to 49%

September 22, 2026
SpaceXAI Releases Grok 4.7: A Larger Base Model at the Same / Price as Grok 4.6
AI & Technology

SpaceXAI Releases Grok 4.7: A Larger Base Model at the Same $2/$6 Price as Grok 4.6

September 22, 2026
Why It’s Important To Unplug Your PC During A Power Outage
AI & Technology

Why It’s Important To Unplug Your PC During A Power Outage

September 22, 2026
Next Post
Yemen’s Houthis Say They Launched Multiple Ballistic Missiles at Israel’s Eilat, Targeted U.S. Ship in Red Sea – Israel News

Yemen's Houthis Say They Launched Multiple Ballistic Missiles at Israel's Eilat, Targeted U.S. Ship in Red Sea - Israel News

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Stay Tuned NOW Streaming Behind The Scenes! – Sept 01

Stay Tuned NOW Streaming Behind The Scenes! – Sept 01

September 19, 2026
U.S. strikes Iran in first military action in weeks

U.S. strikes Iran in first military action in weeks

September 20, 2026
9/11 hero known as ‘man in the red bandana’ awarded Presidential Medal of Freedom

9/11 hero known as ‘man in the red bandana’ awarded Presidential Medal of Freedom

September 15, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!