• bitcoinBitcoin(BTC)$75,520.00-4.12%
  • ethereumEthereum(ETH)$2,395.24-5.96%
  • tetherTether(USDT)$1.00-0.04%
  • binancecoinBNB(BNB)$712.04-1.58%
  • rippleXRP(XRP)$1.28-11.25%
  • usd-coinUSDC(USDC)$1.00-0.02%
  • solanaSolana(SOL)$96.72-6.29%
  • tronTRON(TRX)$0.332433-1.97%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.01-2.76%
  • zcashZcash(ZEC)$1,110.83-5.71%
  • HyperliquidHyperliquid(HYPE)$76.61-4.99%
  • dogecoinDogecoin(DOGE)$0.079789-5.48%
  • RainRain(RAIN)$0.014077-1.79%
  • USDSUSDS(USDS)$1.00-0.03%
  • moneroMonero(XMR)$500.85-2.69%
  • whitebitWhiteBIT Coin(WBT)$77.68-4.80%
  • leo-tokenLEO Token(LEO)$8.88-1.32%
  • chainlinkChainlink(LINK)$10.90-6.34%
  • cardanoCardano(ADA)$0.194708-7.60%
  • stellarStellar(XLM)$0.175335-9.28%
  • Ethena USDeEthena USDe(USDE)$1.00-0.07%
  • daiDai(DAI)$1.000.01%
  • USD1USD1(USD1)$1.00-0.03%
  • bitcoin-cashBitcoin Cash(BCH)$215.08-4.76%
  • litecoinLitecoin(LTC)$51.05-4.28%
  • uniswapUniswap(UNI)$6.24-4.91%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.32-2.66%
  • CantonCanton(CC)$0.090356-7.80%
  • Global DollarGlobal Dollar(USDG)$1.00-0.01%
  • hedera-hashgraphHedera(HBAR)$0.074654-4.31%
  • avalanche-2Avalanche(AVAX)$7.26-4.94%
  • nearNEAR Protocol(NEAR)$2.31-7.66%
  • shiba-inuShiba Inu(SHIB)$0.000005-7.23%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.04%
  • suiSui(SUI)$0.68-6.91%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • tether-goldTether Gold(XAUT)$4,291.88-0.23%
  • crypto-com-chainCronos(CRO)$0.054996-7.74%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • MemeCoreMemeCore(M)$1.121.90%
  • BittensorBittensor(TAO)$217.47-7.33%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • okbOKB(OKB)$109.68-3.42%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.150.03%
  • BitwayBitway(BTW)$0.7017.29%
  • aaveAave(AAVE)$121.24-6.32%
  • pax-goldPAX Gold(PAXG)$4,293.96-0.27%
  • AsterAster(ASTER)$0.68-4.01%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.056904-1.29%
  • mantleMantle(MNT)$0.54-6.05%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Researchers at CMU Introduce TriForce: A Hierarchical Speculative Decoding AI System that is Scalable to Long Sequence Generation

April 20, 2024
in AI & Technology
Reading Time: 5 mins read
A A
Researchers at CMU Introduce TriForce: A Hierarchical Speculative Decoding AI System that is Scalable to Long Sequence Generation
ShareShareShareShareShare

With the widespread deployment of large language models (LLMs) for long content generation, there’s a growing need for efficient long-sequence inference support. However, the key-value (KV) cache, crucial for avoiding re-computation, has become a critical bottleneck, increasing in size linearly with sequence length. The auto-regressive nature of LLMs necessitates loading the entire KV cache for each generated token, leading to low computational core utilization and high latency. While compression methods have been proposed, they often compromise generation quality. LLMs like GPT-4, Gemini, and LWM are gaining prominence in applications like chatbots, vision generation, and financial analysis. However, serving these LLMs efficiently remains challenging due to the auto-regressive nature and the growing memory footprint of the KV cache.

Prior methodologies propose KV cache eviction strategies to reduce the memory footprint of the KV cache, selectively discarding pairs based on eviction policies. This allows models to operate within a limited cache budget. However, such strategies face challenges due to potential information loss, leading to issues like hallucination and contextual incoherency, particularly in long contexts. Speculative decoding, which involves using a lightweight draft model to predict the next tokens, has been introduced to accelerate LLM inference while preserving model output. However, deploying this for long sequence generation presents challenges, including the need for substantial computation to train draft models and the risk of poor speculating performance, especially with existing training-free methods like KV cache eviction strategies.

Researchers from Carnegie Mellon University and Meta AI (FAIR) Introduce TriForce, a hierarchical speculative decoding system designed for scalable long sequence generation. TriForce utilizes the original model weights and dynamic sparse KV cache via retrieval as a draft model, serving as an intermediate layer in the hierarchy. Maintaining the full cache allows for superior KV cache selection using retrieval-based drafting, characterized as lossless compared to eviction-based methods like StreamingLLM and H2O. The hierarchical system addresses dual memory bottlenecks, pairing a lightweight model with a StreamingLLM cache for initial speculations to reduce drafting latency and accelerate end-to-end inference.

TriForce introduces a hierarchical speculative decoding system with retrieval-based KV cache selection. The hierarchical system addresses dual bottlenecks, enhancing speed-up. Retrieval-based drafting segments the KV cache, highlighting relevant information. Lightweight models with StreamingLLM cache accelerate initial speculations, reducing drafting latency. TriForce utilizes model weights and KV cache to enhance LLM inference speed for long sequences. The implementation utilizes Transformers, FlashAttention, and PyTorch CUDA graphs, maintaining full layer sparsity while minimizing kernel launching overhead. 

TriForce evaluation reveals significant speedups, up to 2.31× with a 4K KV cache for Llama2-7B128K on-chip. Offloading to consumer GPUs achieves remarkable efficiency, particularly with Llama2-13B-128K on two RTX 4090 GPUs, 7.94× faster than optimized systems. Llama2-7B-128K with TriForce operates at 0.108s/token, half as slow as auto-regressive baselines on A100. Batched inference also benefits, achieving 1.9× speedup for a batch size of six, each with 19K contexts.

To conclude, this work introduces TriForce, a hierarchical speculative decoding system targeting the efficient serving of LLMs in long contexts. TriForce addresses dual bottlenecks of KV cache and model weights, yielding significant speedups, including up to 2.31× on A100 GPUs and an extraordinary 7.78× on RTX 4090 GPUs. TriForce achieves 0.108s/token, half as slow as auto-regressive baselines on A100. Compared to DeepSpeed-Zero-Inference, TriForce on a single RTX 4090 GPU is 4.86× faster and attains a 1.9× speedup with large batches, showcasing its potential for revolutionizing long-context model serving.


Check out the Paper. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter. Join our Telegram Channel, Discord Channel, and LinkedIn Group.

If you like our work, you will love our newsletter..

Don’t Forget to join our 40k+ ML SubReddit


For Content Partnership, Please Fill Out This Form Here..


YOU MAY ALSO LIKE

Google Releases Gemini 3.8 Live and 3.8 Live Extended Thinking for Production Grade Voice Agents

How To Improve The Audio Quality On Your iPhone

Asjad is an intern consultant at Marktechpost. He is persuing B.Tech in mechanical engineering at the Indian Institute of Technology, Kharagpur. Asjad is a Machine learning and deep learning enthusiast who is always researching the applications of machine learning in healthcare.


🐝 Join the Fastest Growing AI Research Newsletter Read by Researchers from Google + NVIDIA + Meta + Stanford + MIT + Microsoft and many others…


Credit: Source link

ShareTweetSendSharePin

Related Posts

Google Releases Gemini 3.8 Live and 3.8 Live Extended Thinking for Production Grade Voice Agents
AI & Technology

Google Releases Gemini 3.8 Live and 3.8 Live Extended Thinking for Production Grade Voice Agents

September 15, 2026
How To Improve The Audio Quality On Your iPhone
AI & Technology

How To Improve The Audio Quality On Your iPhone

September 15, 2026
Ferrovalle Taps INFORM for AI Smart Yard at Mexico City Rail Hub – Unite.AI
AI & Technology

Ferrovalle Taps INFORM for AI Smart Yard at Mexico City Rail Hub – Unite.AI

September 15, 2026
Google’s Latest Pixel Drop Will Keep You More Connected To Your VIPs
AI & Technology

Google’s Latest Pixel Drop Will Keep You More Connected To Your VIPs

September 15, 2026
Next Post
What it’s actually like to drive this luxury car – reviewing the Bentley Continental GT V8 S

What it's actually like to drive this luxury car – reviewing the Bentley Continental GT V8 S

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Harvey Secures 0M in Fresh Funding, Valuation Climbs to .5B – Unite.AI

Harvey Secures $550M in Fresh Funding, Valuation Climbs to $15.5B – Unite.AI

September 9, 2026
Anthropic Details Disrupted Claude Misuse Across Seven Harm Areas – Unite.AI

Anthropic Details Disrupted Claude Misuse Across Seven Harm Areas – Unite.AI

September 10, 2026
Focus on the Bell Curve

Focus on the Bell Curve

September 11, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!