• bitcoinBitcoin(BTC)$80,290.00-1.21%
  • ethereumEthereum(ETH)$2,570.42-2.68%
  • tetherTether(USDT)$1.00-0.01%
  • binancecoinBNB(BNB)$748.83-2.24%
  • rippleXRP(XRP)$1.38-3.27%
  • usd-coinUSDC(USDC)$1.00-0.01%
  • solanaSolana(SOL)$108.05-3.49%
  • tronTRON(TRX)$0.3423791.43%
  • zcashZcash(ZEC)$1,438.02-6.59%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.02-1.32%
  • HyperliquidHyperliquid(HYPE)$90.81-1.27%
  • dogecoinDogecoin(DOGE)$0.084722-3.74%
  • moneroMonero(XMR)$522.18-10.98%
  • whitebitWhiteBIT Coin(WBT)$81.68-1.77%
  • USDSUSDS(USDS)$1.00-0.01%
  • RainRain(RAIN)$0.013174-5.56%
  • chainlinkChainlink(LINK)$11.99-4.47%
  • cardanoCardano(ADA)$0.219447-2.81%
  • leo-tokenLEO Token(LEO)$8.940.59%
  • stellarStellar(XLM)$0.188990-2.30%
  • uniswapUniswap(UNI)$8.78-3.05%
  • bitcoin-cashBitcoin Cash(BCH)$245.36-2.12%
  • Ethena USDeEthena USDe(USDE)$1.00-0.01%
  • nearNEAR Protocol(NEAR)$3.57-2.57%
  • daiDai(DAI)$1.000.00%
  • litecoinLitecoin(LTC)$56.94-0.54%
  • USD1USD1(USD1)$1.00-0.03%
  • avalanche-2Avalanche(AVAX)$9.765.45%
  • CantonCanton(CC)$0.103553-6.19%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.37-0.01%
  • hedera-hashgraphHedera(HBAR)$0.080241-0.53%
  • suiSui(SUI)$0.81-5.57%
  • MemeCoreMemeCore(M)$1.4410.82%
  • Global DollarGlobal Dollar(USDG)$1.00-0.01%
  • shiba-inuShiba Inu(SHIB)$0.000005-2.33%
  • crypto-com-chainCronos(CRO)$0.057703-3.43%
  • BittensorBittensor(TAO)$249.66-7.35%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.02%
  • tether-goldTether Gold(XAUT)$4,369.57-0.09%
  • Circle USYCCircle USYC(USYC)$1.140.00%
  • okbOKB(OKB)$115.43-2.99%
  • Ripple USDRipple USD(RLUSD)$1.00-0.02%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.15-0.05%
  • aaveAave(AAVE)$135.61-5.87%
  • EthenaEthena(ENA)$0.2030876.07%
  • OndoOndo(ONDO)$0.4101700.69%
  • BitwayBitway(BTW)$0.7322.38%
  • AsterAster(ASTER)$0.73-4.34%
  • mantleMantle(MNT)$0.59-3.16%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

MInference (Milliontokens Inference): A Training-Free Efficient Method for the Pre-Filling Stage of Long-Context LLMs Based on Dynamic Sparse Attention

July 7, 2024
in AI & Technology
Reading Time: 4 mins read
A A
MInference (Milliontokens Inference): A Training-Free Efficient Method for the Pre-Filling Stage of Long-Context LLMs Based on Dynamic Sparse Attention
ShareShareShareShareShare

The computational demands of LLMs, particularly with long prompts, hinder their practical use due to the quadratic complexity of the attention mechanism. For instance, processing a one million-token prompt with an eight-billion-parameter LLM on a single A100 GPU takes about 30 minutes for the initial stage. This leads to significant delays before the model starts generating outputs. While existing methods aim to accelerate this process, they often need to improve accuracy and efficiency. Dynamic sparse attention, which adapts to varying input patterns, can reduce this latency without extensive retraining, unlike fixed sparse methods like Longformer and BigBird.

Researchers from Microsoft Corporation and the University of Surrey have developed MInference (Million-tokens Inference), a method to speed up long-sequence processing in LLMs. By identifying three distinct attention patterns—A-shape, Vertical-Slash, and Block-Sparse—they optimize sparse calculations for GPUs. MInference dynamically builds sparse indices for these patterns during inference, significantly reducing latency without altering pre-training or needing fine-tuning. Tests on various LLMs and benchmarks, such as LLaMA-3-8B-1M and InfiniteBench, show up to a 10x speedup, cutting the pre-filling stage from 30 minutes to 3 minutes on a single A100 GPU while maintaining accuracy. 

YOU MAY ALSO LIKE

Trump Opposes AI Guardrails Amid Chip Selloff

Anthropic’s Claude Takes Bigger Role in Building AI

Sparse attention methods aim to improve Transformer efficiency by reducing the quadratic complexity of attention. These methods include static sparse patterns (e.g., sliding windows, dilated attention), cluster-based approaches (e.g., hash-based, kNN-based), and dynamic sparse attention. However, they typically require pre-training, limiting their direct applicability to ready-to-use LLMs. Recent approaches extend LLM context windows through staged pre-training, modified position embeddings, and external memory modules but do not reduce high inference costs. Other studies optimize pre-filling and decoding in long-context LLMs yet often involve training from scratch or substantial overhead, making them impractical for existing pre-trained models.

Attention weights in long-context LLMs are inherently sparse and dynamic. For instance, in a 128k context, retaining just the top 4k columns covers 96.8% of the total attention. However, the specific tokens attended to by each head can vary greatly with different prompts, making the attention patterns highly context-dependent. Despite this variability, these patterns often exhibit consistent structures across different layers and heads, such as A-shape, Vertical-Slash, and Block-Sparse. Leveraging these patterns, we can significantly optimize sparse computations on GPUs, reducing the computational overhead while maintaining accuracy in long-context LLMs.

The experiments conducted aim to evaluate the effectiveness and efficiency of MInference across multiple benchmarks, including InfiniteBench, RULER, and the Needle in a Haystack task, covering diverse long-context scenarios such as QA, summarization, and retrieval. Four state-of-the-art long-context language models were utilized, including LLaMA-3 and GLM-4, with greedy decoding for consistency. MInference’s performance was tested on various context lengths, demonstrating superiority in maintaining context and processing speed over competing methods. It integrates efficiently with KV cache compression techniques and significantly reduces latency, proving its practical value in optimizing long-context language model performance.

The study tackles the high computational cost and significant latency in the pre-filling stage of long-context LLMs’ attention calculations by introducing MInference. MInference speeds up this process using dynamic sparse attention with specific spatial aggregation patterns: A-shape, Vertical-Slash, and Block-Sparse. A kernel-aware method optimizes the sparse pattern for each attention head, followed by a rapid approximation to create dynamic sparse masks for different inputs, facilitating efficient sparse attention. Testing on benchmarks like InfiniteBench and RULER shows MInference maintains long-context performance while achieving up to a 10x speedup, drastically cutting latency on a single A100 GPU from 30 minutes to 3 minutes for prompts up to 1 million tokens. Similar patterns have potential in multi-modal and encoder-decoder LLMs, indicating promising pre-filling stage acceleration applications.


Check out the Paper, GitHub, and Demo. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter. 

Join our Telegram Channel and LinkedIn Group.

If you like our work, you will love our newsletter..

Don’t Forget to join our 46k+ ML SubReddit


Sana Hassan, a consulting intern at Marktechpost and dual-degree student at IIT Madras, is passionate about applying technology and AI to address real-world challenges. With a keen interest in solving practical problems, he brings a fresh perspective to the intersection of AI and real-life solutions.

🐝 Join the Fastest Growing AI Research Newsletter Read by Researchers from Google + NVIDIA + Meta + Stanford + MIT + Microsoft and many others…


Credit: Source link

ShareTweetSendSharePin

Related Posts

Trump Opposes AI Guardrails Amid Chip Selloff
AI & Technology

Trump Opposes AI Guardrails Amid Chip Selloff

September 20, 2026
Anthropic’s Claude Takes Bigger Role in Building AI
AI & Technology

Anthropic’s Claude Takes Bigger Role in Building AI

September 20, 2026
Anthropic’s Existential Risk Warnings Hijack Larger AI Debate
AI & Technology

Anthropic’s Existential Risk Warnings Hijack Larger AI Debate

September 20, 2026
Anthropic Investor Franklin: AI Safety Concerns Won’t Slow Spending
AI & Technology

Anthropic Investor Franklin: AI Safety Concerns Won’t Slow Spending

September 20, 2026
Next Post
Severe weather causes flooding across the Midwest

Severe weather causes flooding across the Midwest

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
US 10-year Treasury yield briefly surges past 5% on sky-high diesel prices

US 10-year Treasury yield briefly surges past 5% on sky-high diesel prices

September 14, 2026
New video of deadly cargo plane crash in Miami

New video of deadly cargo plane crash in Miami

September 14, 2026
Hierarchical NeRF with JAX3D for Volumetric Rendering, Novel-View Synthesis, and 3D Reconstruction

Hierarchical NeRF with JAX3D for Volumetric Rendering, Novel-View Synthesis, and 3D Reconstruction

September 13, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!