• bitcoinBitcoin(BTC)$81,179.005.96%
  • ethereumEthereum(ETH)$2,612.656.53%
  • tetherTether(USDT)$1.000.05%
  • binancecoinBNB(BNB)$761.672.61%
  • rippleXRP(XRP)$1.417.89%
  • usd-coinUSDC(USDC)$1.000.02%
  • solanaSolana(SOL)$113.4510.95%
  • tronTRON(TRX)$0.3381360.95%
  • zcashZcash(ZEC)$1,568.206.40%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.030.41%
  • HyperliquidHyperliquid(HYPE)$93.057.66%
  • dogecoinDogecoin(DOGE)$0.0876736.66%
  • moneroMonero(XMR)$582.5712.59%
  • whitebitWhiteBIT Coin(WBT)$83.195.41%
  • USDSUSDS(USDS)$1.000.03%
  • RainRain(RAIN)$0.0133855.56%
  • chainlinkChainlink(LINK)$12.327.39%
  • cardanoCardano(ADA)$0.2266919.32%
  • leo-tokenLEO Token(LEO)$8.900.06%
  • stellarStellar(XLM)$0.1931245.10%
  • uniswapUniswap(UNI)$8.9614.50%
  • bitcoin-cashBitcoin Cash(BCH)$256.158.87%
  • nearNEAR Protocol(NEAR)$3.8619.33%
  • Ethena USDeEthena USDe(USDE)$1.000.05%
  • daiDai(DAI)$1.00-0.01%
  • litecoinLitecoin(LTC)$58.427.34%
  • CantonCanton(CC)$0.1112095.10%
  • USD1USD1(USD1)$1.000.06%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.361.73%
  • avalanche-2Avalanche(AVAX)$8.5211.20%
  • hedera-hashgraphHedera(HBAR)$0.0790405.32%
  • suiSui(SUI)$0.828.13%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • shiba-inuShiba Inu(SHIB)$0.0000053.95%
  • MemeCoreMemeCore(M)$1.315.05%
  • crypto-com-chainCronos(CRO)$0.0598332.84%
  • BittensorBittensor(TAO)$251.887.01%
  • paypal-usdPayPal USD(PYUSD)$1.000.04%
  • tether-goldTether Gold(XAUT)$4,371.000.55%
  • Circle USYCCircle USYC(USYC)$1.140.03%
  • okbOKB(OKB)$115.953.46%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.14-0.02%
  • aaveAave(AAVE)$141.328.92%
  • AsterAster(ASTER)$0.774.01%
  • mantleMantle(MNT)$0.628.96%
  • Pump.funPump.fun(PUMP)$0.0042174.83%
  • OndoOndo(ONDO)$0.3994643.20%
  • polkadotPolkadot(DOT)$1.132.48%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

FastGen: Cutting GPU Memory Costs Without Compromising on LLM Quality

May 13, 2024
in AI & Technology
Reading Time: 4 mins read
A A
FastGen: Cutting GPU Memory Costs Without Compromising on LLM Quality
ShareShareShareShareShare

Autoregressive language models (ALMs) have proven their capability in machine translation, text generation, etc. However, these models pose challenges, including computational complexity and GPU memory usage. Despite great success in various applications, there is an urgent need to find a cost-effective way to serve these models. Moreover, the generative inference of large language models (LLMs) utilizes the KV Cache mechanism to enhance the generation speed. Still, an increase in model size and generation length leads to an increase in memory usage of the KV cache. When memory usage exceeds GPU capacity, the generative inference of LLMs resorts to offloading.

Many works have been carried out to enhance the model efficiency for LLMs, e.g., one such method is to skip multiple tokens at a particular time stamp. Recently, a technique that adds a token selection task to the original BERT model learns to select performance-crucial tokens and detect unimportant tokens to prune using a designed learnable threshold. However, these models are only applied to non-autoregressive models and require an extra re-training phrase, making them less suitable for auto-regressive LLMs like ChatGPT and Llama. It is important to consider pruning tokens’ potential within the KV cache of auto-regressive LLMs to fill this gap.   

Researchers from the University of Illinois Urbana-Champaign and Microsoft proposed FastGen, a highly effective technique to enhance the inference efficiency of LLMs without any loss in visible quality, using lightweight model profiling and adaptive key-value caching. FastGen evicts long-range contexts on attention heads by the KV cache construction in an adaptive manner. Moreover, it is deployed using lightweight attention profiling, which has been used to guide the construction of the adaptive KV cache without resource-intensive fine-tuning or re-training. FastGen is capable of reducing GPU memory usage with negligible generation quality loss.

The adaptive KV Cache compression introduced by the researchers reduces the memory footprint of generative inference for LLMs. In this method, there are two steps for a generative model inference which are involved:

  • Prompt Encoding: The attention module needs to collect contextual information from all the preceding i-1 tokens for the i-th token generated by autoregressive transformer-based LLM.
  • Token Generation: When prompt encoding is completed, LLM generates the output token by token, and for each step, the new token(s) generated in the previous step are encoded using the LLM. 

For 30B models, FastGen outperforms all non-adaptive KV compression methods and achieves a higher KV cache reduction ratio with an increase in model size, keeping the model’s quality unaffected. For example, FastGen gets a 44.9% pruned ratio on Llama 1-65B, compared to a 16.9% pruned ratio on Llama 1-7B, achieving a 45% win rate. Further, sensitivity analysis was performed on FastGen by choosing different hyper-parameters. Since the model maintains a win rate of 45%, the study shows no visible impact on generation quality after changing the hyper-parameter.   

In conclusion, researchers from the University of Illinois Urbana-Champaign and Microsoft proposed FastGen, a new technique to enhance LLMs inference efficiency with no loss in visible quality, using lightweight model profiling and adaptive key-value caching. Also, the adaptive KV Cache compression introduced by researchers is constructed using FastGen to reduce the memory footprint of generative inference for LLMs. Future work includes integrating FastGen with other model compression approaches, e.g., quantization and distillation, grouped-query attention, etc.


Check out the Paper. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter. Join our Telegram Channel, Discord Channel, and LinkedIn Group.

If you like our work, you will love our newsletter..

Don’t Forget to join our 42k+ ML SubReddit


YOU MAY ALSO LIKE

AI Almost Led The US Military To Start A War With China, Report Says

Jina AI Releases jina-ocr-v1: A 3.4B MoE Document Parser With Built-In Speculative Decoding for Low-Budget GPUs

Sajjad Ansari is a final year undergraduate from IIT Kharagpur. As a Tech enthusiast, he delves into the practical applications of AI with a focus on understanding the impact of AI technologies and their real-world implications. He aims to articulate complex AI concepts in a clear and accessible manner.


[Recommended Read] Rightsify’s GCX: Your Go-To Source for High-Quality, Ethically Sourced, Copyright-Cleared AI Music Training Datasets with Rich Metadata


Credit: Source link

ShareTweetSendSharePin

Related Posts

AI Almost Led The US Military To Start A War With China, Report Says
AI & Technology

AI Almost Led The US Military To Start A War With China, Report Says

September 18, 2026
Jina AI Releases jina-ocr-v1: A 3.4B MoE Document Parser With Built-In Speculative Decoding for Low-Budget GPUs
AI & Technology

Jina AI Releases jina-ocr-v1: A 3.4B MoE Document Parser With Built-In Speculative Decoding for Low-Budget GPUs

September 18, 2026
Sony Music And UMG Say Suno’s New Models Still Violates Their Copyright
AI & Technology

Sony Music And UMG Say Suno’s New Models Still Violates Their Copyright

September 18, 2026
Anthropic Taps Accenture’s Faculty for Embedded AI Model Evaluation – Unite.AI
AI & Technology

Anthropic Taps Accenture’s Faculty for Embedded AI Model Evaluation – Unite.AI

September 18, 2026
Next Post
Senate advances aid bill for Ukraine, Israel and Taiwan without provisions for U.S. border

Senate advances aid bill for Ukraine, Israel and Taiwan without provisions for U.S. border

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Trump names Adam Telle as acting Army secretary

Trump names Adam Telle as acting Army secretary

September 17, 2026
Raskin says its more important for Democrats to address healthcare

Raskin says its more important for Democrats to address healthcare

September 16, 2026
County coroner says Pennsylvania infant died from measles

County coroner says Pennsylvania infant died from measles

September 17, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!