• bitcoinBitcoin(BTC)$84,439.000.61%
  • ethereumEthereum(ETH)$2,697.070.34%
  • tetherTether(USDT)$1.00-0.01%
  • binancecoinBNB(BNB)$771.77-0.28%
  • rippleXRP(XRP)$1.51-2.72%
  • usd-coinUSDC(USDC)$1.00-0.01%
  • solanaSolana(SOL)$120.540.03%
  • tronTRON(TRX)$0.333110-1.26%
  • zcashZcash(ZEC)$1,638.816.90%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.063.41%
  • HyperliquidHyperliquid(HYPE)$93.772.12%
  • dogecoinDogecoin(DOGE)$0.095999-1.91%
  • moneroMonero(XMR)$562.090.85%
  • chainlinkChainlink(LINK)$14.110.03%
  • whitebitWhiteBIT Coin(WBT)$84.240.55%
  • USDSUSDS(USDS)$1.000.01%
  • cardanoCardano(ADA)$0.251660-1.81%
  • RainRain(RAIN)$0.01277715.11%
  • leo-tokenLEO Token(LEO)$9.041.22%
  • stellarStellar(XLM)$0.214079-1.96%
  • bitcoin-cashBitcoin Cash(BCH)$331.78-1.68%
  • nearNEAR Protocol(NEAR)$5.022.81%
  • uniswapUniswap(UNI)$9.741.47%
  • litecoinLitecoin(LTC)$71.51-0.71%
  • CantonCanton(CC)$0.1373914.93%
  • Ethena USDeEthena USDe(USDE)$1.00-0.02%
  • suiSui(SUI)$1.17-0.15%
  • avalanche-2Avalanche(AVAX)$10.740.96%
  • daiDai(DAI)$1.000.00%
  • USD1USD1(USD1)$1.000.01%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.567.18%
  • hedera-hashgraphHedera(HBAR)$0.092992-1.27%
  • BittensorBittensor(TAO)$320.042.78%
  • shiba-inuShiba Inu(SHIB)$0.000006-0.93%
  • crypto-com-chainCronos(CRO)$0.0663170.59%
  • Global DollarGlobal Dollar(USDG)$1.00-0.02%
  • MemeCoreMemeCore(M)$1.230.58%
  • paypal-usdPayPal USD(PYUSD)$1.000.00%
  • BitwayBitway(BTW)$1.0211.90%
  • EthenaEthena(ENA)$0.268105-0.91%
  • tether-goldTether Gold(XAUT)$4,277.63-0.11%
  • OndoOndo(ONDO)$0.53-2.19%
  • quant-networkQuant(QNT)$177.0278.30%
  • okbOKB(OKB)$120.91-0.50%
  • Ripple USDRipple USD(RLUSD)$1.00-0.01%
  • Circle USYCCircle USYC(USYC)$1.140.00%
  • aaveAave(AAVE)$154.970.62%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.15-0.05%
  • mantleMantle(MNT)$0.680.08%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

CPU-GPU I/O-Aware LLM Inference Reduces Latency in GPUs by Optimizing CPU-GPU Interactions

December 7, 2024
in AI & Technology
Reading Time: 7 mins read
A A
CPU-GPU I/O-Aware LLM Inference Reduces Latency in GPUs by Optimizing CPU-GPU Interactions
ShareShareShareShareShare

LLMs are driving major advances in research and development today. A significant shift has been observed in research objectives and methodologies toward an LLM-centric approach. However, they are associated with high expenses, making LLMs for large-scale utilization inaccessible to many. It is, therefore, a significant challenge to reduce the latency of operations, especially in dynamic applications that demand responsiveness.

KV cache is used for autoregressive decoding in LLMs. It stores key-value pairs in multi-headed attention during the pre-filling phase of inference. During the decoding stage, new KV pairs get appended to the memory. KV cache stores the intermediate key and value activations in the attention mechanism, thus reducing complexity from quadratic to linear order. KV cache allows for improved efficiency but grows linearly with batch size, sequence length, and model size. The growing memory size of the KV cache exceeds the handling capacity of GPUs, and transferring it to the CPU introduces several bottlenecks, increasing latency while reducing throughput.

YOU MAY ALSO LIKE

Why We Won’t Know How Visible The iPhone Duo’s Crease Is For A Long Time

How Powerful Of A Power Bank Do You Need To Safely Charge A Laptop?

PCIe interfaces become a limiting factor, especially when transferring the cache from the CPU to the GPU for computation. Slow PCIe interfaces can result in latency exceeding normal levels by an order of magnitude, leading to substantial GPU idle time.

Previous work has attempted to mitigate the issue of slow PCIe performance. Still, these approaches often fail due to mismatched data transfer and GPU computation times, particularly with large batch and context sizes. Others depended on CPU resources, which again became a limiting factor. This article discusses a novel approach to PCIe and GPU optimization.

University of Southern California researchers propose an efficient CPU-GPU I/O-aware LLM inference method for optimized PCIe utilization. It leverages partial KV cache recomputation and asynchronous overlapping to address the system bottleneck of loading large KV caches. Their process involves transferring smaller activation segments of the cache to the GPU rather than transferring the entire KV cache. The GPU then reconstructs the whole cache memory from these smaller activation bits. The key lies in computing attention scores that ensure minimal information loss.

The authors propose a fully automated method for determining recomputation and communication splits. This work consists of three modules to minimize GPU latency:

  1. Profiler Module: Collects system hardware information, such as PCIe bandwidth and GPU processing speed.
  2. Scheduler Module: Formulates the problem as a linear programming task to determine the optimal KV split point using hardware information and user configuration. The objective is to maximize the overlap between computation and communication processes.
  3. Runtime Module: Coordinates data transfer between the two devices and manages memory allocations.

The Scheduler Module, which is responsible for finding the optimal KV split, works in two ways:

Row-by-Row Schedule: Reduces latency with a row-by-row execution plan. Here, the GPU begins reconstructing the KV cache while the remaining activations are asynchronously loading. Column-by-Column Schedule: Maximizes throughput and accommodates significant batch size inference by reusing model weights across batches. It overlaps the transmission of KV cache and activations with the computation of MHA (multi-headed attention) across multiple batches instead of processing each layer sequentially in a batch.Further using a six-process communication parallelism strategy, the Runtime Module enables concurrent GPU computation and CPU-GPU communication.

The authors tested the proposed framework for efficient LLM inference using an NVIDIA A100 GPU connected to a CPU via a PCIe 4.0 x16 interface. Experiments were conducted with two objectives to assess the framework’s performance:

  • Latency-Oriented Workload: The proposed method outperformed baselines, reducing latency by 35.8%.
  • Throughput-Oriented Workload: The method achieved up to a 29% improvement relative to the baseline.

Conclusion:

The CPU-GPU I/O-aware LLM inference method efficiently reduces latency while increasing throughput in LLM inference. It leverages partial KV cache recomputation and overlaps it with data transmission to minimize idle GPU time and enhance efficiency.


Check out the Paper. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter and join our Telegram Channel and LinkedIn Group. If you like our work, you will love our newsletter.. Don’t Forget to join our 60k+ ML SubReddit.

🚨 [Partner with us]: ‘Next Magazine/Report- Open Source AI in Production’


Adeeba Alam Ansari is currently pursuing her Dual Degree at the Indian Institute of Technology (IIT) Kharagpur, earning a B.Tech in Industrial Engineering and an M.Tech in Financial Engineering. With a keen interest in machine learning and artificial intelligence, she is an avid reader and an inquisitive individual. Adeeba firmly believes in the power of technology to empower society and promote welfare through innovative solutions driven by empathy and a deep understanding of real-world challenges.

🚨🚨FREE AI WEBINAR: ‘Fast-Track Your LLM Apps with deepset & Haystack'(Promoted)


Credit: Source link

ShareTweetSendSharePin

Related Posts

Why We Won’t Know How Visible The iPhone Duo’s Crease Is For A Long Time
AI & Technology

Why We Won’t Know How Visible The iPhone Duo’s Crease Is For A Long Time

September 27, 2026
How Powerful Of A Power Bank Do You Need To Safely Charge A Laptop?
AI & Technology

How Powerful Of A Power Bank Do You Need To Safely Charge A Laptop?

September 27, 2026
Your Old GPU Could Be Worth More Than You Think
AI & Technology

Your Old GPU Could Be Worth More Than You Think

September 26, 2026
Sarvam AI Releases Saaras V4: A Speech-to-Text Model for All 22 Indian Languages and Global English
AI & Technology

Sarvam AI Releases Saaras V4: A Speech-to-Text Model for All 22 Indian Languages and Global English

September 26, 2026
Next Post
This AI Paper from UCLA Unveils ‘2-Factor Retrieval’ for Revolutionizing Human-AI Decision-Making in Radiology

This AI Paper from UCLA Unveils '2-Factor Retrieval' for Revolutionizing Human-AI Decision-Making in Radiology

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Emerging Market Debt: The Next Frontier For AI Disruption?

Emerging Market Debt: The Next Frontier For AI Disruption?

September 26, 2026
Feds scrutinize  billion in unusual Kalshi trades amid ‘wash trading’ concerns

Feds scrutinize $5 billion in unusual Kalshi trades amid ‘wash trading’ concerns

September 23, 2026
Sen. Tim Scott says Darline Graham is ‘ready for primetime’

Sen. Tim Scott says Darline Graham is ‘ready for primetime’

September 25, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!