• bitcoinBitcoin(BTC)$81,354.000.29%
  • ethereumEthereum(ETH)$2,636.860.22%
  • tetherTether(USDT)$1.000.00%
  • binancecoinBNB(BNB)$763.60-0.34%
  • rippleXRP(XRP)$1.431.81%
  • usd-coinUSDC(USDC)$1.000.00%
  • solanaSolana(SOL)$111.38-2.03%
  • tronTRON(TRX)$0.3393520.31%
  • zcashZcash(ZEC)$1,483.730.43%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.00-2.77%
  • HyperliquidHyperliquid(HYPE)$92.520.67%
  • dogecoinDogecoin(DOGE)$0.0896742.18%
  • moneroMonero(XMR)$552.60-1.01%
  • RainRain(RAIN)$0.0139333.08%
  • whitebitWhiteBIT Coin(WBT)$83.02-0.37%
  • USDSUSDS(USDS)$1.00-0.03%
  • chainlinkChainlink(LINK)$12.542.06%
  • cardanoCardano(ADA)$0.2307283.93%
  • leo-tokenLEO Token(LEO)$8.920.42%
  • stellarStellar(XLM)$0.1996683.04%
  • uniswapUniswap(UNI)$8.70-2.90%
  • bitcoin-cashBitcoin Cash(BCH)$255.070.52%
  • Ethena USDeEthena USDe(USDE)$1.00-0.02%
  • nearNEAR Protocol(NEAR)$3.53-5.22%
  • daiDai(DAI)$1.000.00%
  • litecoinLitecoin(LTC)$57.861.73%
  • CantonCanton(CC)$0.1114311.37%
  • USD1USD1(USD1)$1.000.00%
  • avalanche-2Avalanche(AVAX)$9.7618.23%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.391.72%
  • hedera-hashgraphHedera(HBAR)$0.0821643.92%
  • suiSui(SUI)$0.878.02%
  • shiba-inuShiba Inu(SHIB)$0.0000062.62%
  • MemeCoreMemeCore(M)$1.438.98%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • BittensorBittensor(TAO)$263.394.89%
  • crypto-com-chainCronos(CRO)$0.059657-0.16%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.02%
  • tether-goldTether Gold(XAUT)$4,373.63-0.09%
  • Circle USYCCircle USYC(USYC)$1.140.00%
  • okbOKB(OKB)$118.361.48%
  • Ripple USDRipple USD(RLUSD)$1.00-0.01%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.15-0.15%
  • aaveAave(AAVE)$142.292.84%
  • OndoOndo(ONDO)$0.4315918.30%
  • AsterAster(ASTER)$0.771.93%
  • mantleMantle(MNT)$0.630.75%
  • EthenaEthena(ENA)$0.20306520.96%
  • Pump.funPump.fun(PUMP)$0.004171-5.31%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

NVIDIA Researchers Introduce KVTC Transform Coding Pipeline to Compress Key-Value Caches by 20x for Efficient LLM Serving

February 11, 2026
in AI & Technology
Reading Time: 7 mins read
A A
NVIDIA Researchers Introduce KVTC Transform Coding Pipeline to Compress Key-Value Caches by 20x for Efficient LLM Serving
ShareShareShareShareShare

Serving Large Language Models (LLMs) at scale is a massive engineering challenge because of Key-Value (KV) cache management. As models grow in size and reasoning capability, the KV cache footprint increases and becomes a major bottleneck for throughput and latency. For modern Transformers, this cache can occupy multiple gigabytes.

NVIDIA researchers have introduced KVTC (KV Cache Transform Coding). This lightweight transform coder compresses KV caches for compact on-GPU and off-GPU storage. It achieves up to 20x compression while maintaining reasoning and long-context accuracy. For specific use cases, it can reach 40x or higher.

YOU MAY ALSO LIKE

Why Is Your iPad Not Charging (And How To Fix It)

How To Block And Unblock A Number On Your Android Phone

https://arxiv.org/pdf/2511.01815

The Memory Dilemma in LLM Inference

In production, inference frameworks treat local KV caches like databases. Strategies like prefix sharing promote the reuse of caches to speed up responses. However, stale caches consume scarce GPU memory. Developers currently face a difficult choice:

  • Keep the cache: Occupies memory needed for other users.
  • Discard the cache: Incurs the high cost of recomputation.
  • Offload the cache: Moves data to CPU DRAM or SSDs, leading to transfer overheads.

KVTC largely mitigates this dilemma by lowering the cost of on-chip retention and reducing the bandwidth required for offloading.

https://arxiv.org/pdf/2511.01815

How the KVTC Pipeline Works?

The method is inspired by classical media compression. It applies a learned orthonormal transform, followed by adaptive quantization and entropy coding.

1. Feature Decorrelation (PCA)

Different attention heads often show similar patterns and a high degree of correlation. KVTC uses Principal Component Analysis (PCA) to linearly decorrelate features. Unlike other methods that calculate a separate decomposition for every prompt, KVTC computes the PCA basis matrix V once on a calibration dataset. This matrix is then reused for all future caches at inference time.

2. Adaptive Quantization

The system exploits the PCA ordering to allocate a fixed bit budget across coordinates. High-variance components receive more bits, while others receive fewer. KVTC uses a dynamic programming (DP) algorithm to find the optimal bit allocation that minimizes reconstruction error. Crucially, the DP often assigns 0 bits to trailing principal components, allowing for early dimensionality reduction and faster performance.

3. Entropy Coding

The quantized symbols are packed and compressed using the DEFLATE algorithm. To maintain speed, KVTC leverages the nvCOMP library, which enables parallel compression and decompression directly on the GPU.

Protecting Critical Tokens

Not all tokens are compressed equally. KVTC avoids compressing two specific types of tokens because they contribute disproportionately to attention accuracy:

  • Attention Sinks: The 4 oldest tokens in the sequence.
  • Sliding Window: The 128 most recent tokens.

Ablation studies show that compressing these specific tokens can significantly lower or even collapse accuracy at high compression ratios.

Benchmarks and Efficiency

The research team tested KVTC with models like Llama-3.1, Mistral-NeMo, and R1-Qwen-2.5.

  • Accuracy: At 16x compression (roughly 20x after DEFLATE), the model consistently maintains results within 1 score point of vanilla models.
  • TTFT Reduction: For an 8K context length, kvtc can reduce Time-To-First-Token (TTFT) by up to 8x compared to full recomputation.
  • Speed: Calibration is fast; for a 12B model, it can be completed within 10 minutes on an NVIDIA H100 GPU.
  • Storage Overhead: The extra data stored per model is small, representing only 2.4% of model parameters for Llama-3.3-70B.

KVTC is a practical building block for memory-efficient LLM serving. It does not modify model weights and is directly compatible with other token eviction methods.

https://arxiv.org/pdf/2511.01815

Key Takeaways

  • High Compression with Low Accuracy Loss: KVTC achieves a standard 20x compression ratio while maintaining results within 1 score point of vanilla (uncompressed) models across most reasoning and long-context benchmarks.
  • Transform Coding Pipeline: The method utilizes a pipeline inspired by classical media compression, combining PCA-based feature decorrelation, adaptive quantization via dynamic programming, and lossless entropy coding (DEFLATE).
  • Critical Token Protection: To maintain model performance, KVTC avoids compressing the 4 oldest ‘attention sink’ tokens and a ‘sliding window’ of the 128 most recent tokens.
  • Operational Efficiency: The system is ‘tuning-free,’ requiring only a brief initial calibration (under 10 minutes for a 12B model) that leaves model parameters unchanged and adds minimal storage overhead—only 2.4% for a 70B model.
  • Significant Latency Reduction: By reducing the volume of data stored and transferred, KVTC can reduce Time-To-First-Token (TTFT) by up to 8x compared to the full recomputation of KV caches for long contexts.

Check out the Paper here. Also, feel free to follow us on Twitter and don’t forget to join our 100k+ ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well.

The post NVIDIA Researchers Introduce KVTC Transform Coding Pipeline to Compress Key-Value Caches by 20x for Efficient LLM Serving appeared first on MarkTechPost.

Credit: Source link

ShareTweetSendSharePin

Related Posts

Why Is Your iPad Not Charging (And How To Fix It)
AI & Technology

Why Is Your iPad Not Charging (And How To Fix It)

September 19, 2026
How To Block And Unblock A Number On Your Android Phone
AI & Technology

How To Block And Unblock A Number On Your Android Phone

September 19, 2026
Google Gemini Also Escaped Its Testing Environment And Hacked Three Companies
AI & Technology

Google Gemini Also Escaped Its Testing Environment And Hacked Three Companies

September 19, 2026
What Is AI Agent Memory? Short-Term, Long-Term, Episodic, and Semantic Memory Explained – Unite.AI
AI & Technology

What Is AI Agent Memory? Short-Term, Long-Term, Episodic, and Semantic Memory Explained – Unite.AI

September 19, 2026
Next Post
Students come together for proposal surprise

Students come together for proposal surprise

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Trump asks Supreme Court to allow mail ballot restrictions to move forward for third time

Trump asks Supreme Court to allow mail ballot restrictions to move forward for third time

September 16, 2026
Wall St set to open higher as easing oil prices boost Fed hike relief – reuters.com

Wall St set to open higher as easing oil prices boost Fed hike relief – reuters.com

September 17, 2026
May World Oil Production At Post-War Low

May World Oil Production At Post-War Low

September 16, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!