• bitcoinBitcoin(BTC)$77,180.000.10%
  • ethereumEthereum(ETH)$2,520.880.17%
  • tetherTether(USDT)$1.000.00%
  • binancecoinBNB(BNB)$725.790.23%
  • rippleXRP(XRP)$1.360.80%
  • usd-coinUSDC(USDC)$1.000.00%
  • solanaSolana(SOL)$101.60-0.32%
  • tronTRON(TRX)$0.3399400.53%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.00-3.12%
  • zcashZcash(ZEC)$1,122.51-3.23%
  • HyperliquidHyperliquid(HYPE)$79.580.56%
  • dogecoinDogecoin(DOGE)$0.0846110.87%
  • RainRain(RAIN)$0.0157762.03%
  • moneroMonero(XMR)$538.174.32%
  • USDSUSDS(USDS)$1.00-0.01%
  • whitebitWhiteBIT Coin(WBT)$80.230.17%
  • chainlinkChainlink(LINK)$11.49-0.15%
  • leo-tokenLEO Token(LEO)$9.15-0.05%
  • cardanoCardano(ADA)$0.2067140.95%
  • stellarStellar(XLM)$0.1798391.23%
  • Ethena USDeEthena USDe(USDE)$1.00-0.01%
  • daiDai(DAI)$1.00-0.01%
  • bitcoin-cashBitcoin Cash(BCH)$225.36-0.53%
  • USD1USD1(USD1)$1.000.00%
  • litecoinLitecoin(LTC)$53.490.96%
  • uniswapUniswap(UNI)$6.345.77%
  • CantonCanton(CC)$0.0972910.48%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.381.51%
  • Global DollarGlobal Dollar(USDG)$1.00-0.01%
  • avalanche-2Avalanche(AVAX)$7.39-0.46%
  • hedera-hashgraphHedera(HBAR)$0.0745710.64%
  • shiba-inuShiba Inu(SHIB)$0.0000053.32%
  • nearNEAR Protocol(NEAR)$2.35-2.75%
  • suiSui(SUI)$0.720.03%
  • crypto-com-chainCronos(CRO)$0.0594446.01%
  • paypal-usdPayPal USD(PYUSD)$1.000.01%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • MemeCoreMemeCore(M)$1.18-0.63%
  • tether-goldTether Gold(XAUT)$4,349.710.07%
  • Circle USYCCircle USYC(USYC)$1.140.00%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • okbOKB(OKB)$113.951.00%
  • BittensorBittensor(TAO)$232.37-0.86%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.150.17%
  • aaveAave(AAVE)$125.210.95%
  • pax-goldPAX Gold(PAXG)$4,354.210.04%
  • mantleMantle(MNT)$0.57-1.95%
  • AsterAster(ASTER)$0.680.99%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.0574105.25%
  • polkadotPolkadot(DOT)$1.02-1.51%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Seeking Speed without Loss in Large Language Models? Meet EAGLE: A Machine Learning Framework Setting New Standards for Lossless Acceleration

February 2, 2024
in AI & Technology
Reading Time: 5 mins read
A A
Seeking Speed without Loss in Large Language Models? Meet EAGLE: A Machine Learning Framework Setting New Standards for Lossless Acceleration
ShareShareShareShareShare

For LLMs, auto-regressive decoding is now considered the gold standard. Because LLMs generate output tokens individually, the procedure is time-consuming and expensive. Methods based on speculative sampling provide an answer to this problem. In the first, called the “draft” phase, LLMs are hypothesized at little cost; in the second, called the “verification” phase, all of the proposed tokens are checked in parallel using a single forward pass of the LLM. Speed is greatly improved by parallelizing speculative sampling, which allows for producing many post-check tokens for every LLM forward pass.

Speculative sampling aims to find a preliminary model that is comparable to the original LLM in terms of latency but faster. In most cases, a lower-parameter LLM derived from the same data set as the draft model is used in speculative sampling. 

Speeding up speculative sampling requires lowering the time overhead and increasing the draft’s acceptance rate by the original LLM. However, the drafts produced by these systems are less precise, limiting their potential.

Recent studies by Peking University, Microsoft Research, University of Waterloo, and Vector Institute present EAGLE (Extrapolation Algorithm for Greater Language-model Efficiency). It is a straightforward framework that departs from direct token prediction and executes auto-regressive operations at the feature level based on the observation that feature-level auto-regression is easier to handle than token-level auto-regression. EAGLE avoids the uncertainty in feature-level auto-regression by using a token sequence advanced by a one-time step. 

Theoretically, in both the greedy and non-greedy settings, EAGLE is guaranteed to preserve the output distribution and does not involve fine-tuning the original LLM. In certain instances, acceleration could cause LLM outputs to be incorrect or even hazardous, preventing any degradation. Lookahead and Medusa, on the other hand, are solely concerned with greedy situations. Compared to Medusa’s 0.6, EAGLE’s draft accuracy of about 0.8 is significantly better, achieved with a model that only includes a transformer decoder layer.

The study also offers views on aspects contributing to EAGLE’s effectiveness and introduces the simple yet efficient structure. These factors might be of independent relevance to other speculative sampling approaches. The foundation of EAGLE are these two findings: 

  • Top-layer features are more effective than bottom-layer token embeddings with the same lightweight network.
  • Draft models that only input top-layer features are severely limited in performance due to the inherent uncertainty in the sampling process. 

That is why it is critical to incorporate the token representing the sample results into the preliminary model.

The team tested EAGLE on the MT-bench, a realistic benchmark miming real-world scenarios and applications. This benchmark includes multi-turn instructions similar to ChatGPT dialogues. Because it has been used state-of-the-art to exhibit speedup ratios by Lookahead and Medusa, they have also decided to use it. This decision makes it easy to compare the proposed method to these standards impartially and straightforwardly. With a greedy decoding configuration, EAGLE provides a 3x acceleration for Vicuna-13B and LLaMA2-Chat 13B, 70B, which is theoretically certain to preserve the original LLM’s text distribution and is immediately usable. EAGLE outperforms the newly suggested speculative sampling-based frameworks Lookahead and Medusa with a speedup of 2x and a speedup of 1.6x, respectively. With EAGLE, performance is improved, and LLM systems’ throughput is doubled. 

EAGLE runs in tandem with other acceleration or throughput-enhancing techniques like quantization and compilation. The operational expenses of LLM systems could be further reduced by combining EAGLE with these approaches. Using gpt-fast1, EAGLE can increase the throughput of LLaMA2-Chat 7B decoding on a single RTX 3090 GPU from 24.5 to 160.4 tokens/s. Low training expenses are a feature of EAGLE. To train a decoder layer with less than 1 billion parameters for the LLaMA2-Chat 70B model, EAGLE uses the ShareGPT dataset with no more than 70k dialogues. On four A100 (40G) GPUs, the training takes about a day or two to finish. EAGLE can accelerate each query in real-world scenarios with just one training session. The amortized training cost of EAGLE falls to zero as the number of queries rises.


Check out the Paper and GitHub. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter and Google News. Join our 36k+ ML SubReddit, 41k+ Facebook Community, Discord Channel, and LinkedIn Group.

If you like our work, you will love our newsletter..

Don’t Forget to join our Telegram Channel


YOU MAY ALSO LIKE

Blizzard Is Reviving StarCraft As An Open-World Shooter, But It’ll Be A Long Wait

Diablo V Is Coming Out In Spring 2029

Dhanshree Shenwai is a Computer Science Engineer and has a good experience in FinTech companies covering Financial, Cards & Payments and Banking domain with keen interest in applications of AI. She is enthusiastic about exploring new technologies and advancements in today’s evolving world making everyone’s life easy.


🎯 [FREE AI WEBINAR] ‘Using ANN for Vector Search at Speed & Scale (Demo on AWS)’ (Feb 5, 2024)


Credit: Source link

ShareTweetSendSharePin

Related Posts

Blizzard Is Reviving StarCraft As An Open-World Shooter, But It’ll Be A Long Wait
AI & Technology

Blizzard Is Reviving StarCraft As An Open-World Shooter, But It’ll Be A Long Wait

September 12, 2026
Diablo V Is Coming Out In Spring 2029
AI & Technology

Diablo V Is Coming Out In Spring 2029

September 12, 2026
Fly Language Model (FLM) Wires the Full Fruit Fly Connectome Into a Frozen 1.2B LLM, and Its Own Controls Show the Wiring Does Not Help
AI & Technology

Fly Language Model (FLM) Wires the Full Fruit Fly Connectome Into a Frozen 1.2B LLM, and Its Own Controls Show the Wiring Does Not Help

September 12, 2026
Altman Says OpenAI Will Match Anthropic’s Embedded Evaluator Pledge – Unite.AI
AI & Technology

Altman Says OpenAI Will Match Anthropic’s Embedded Evaluator Pledge – Unite.AI

September 12, 2026
Next Post
#amazon drops it’s deal to buy Roomba maker iRobot #technology

#amazon drops it’s deal to buy Roomba maker iRobot #technology

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Disinherit Our Trust Fund Baby?

Disinherit Our Trust Fund Baby?

September 6, 2026
Amazon, Qualcomm Deal Broadens the AI Chip Race | Bloomberg Tech 9/08/2026

Amazon, Qualcomm Deal Broadens the AI Chip Race | Bloomberg Tech 9/08/2026

September 12, 2026
Should You Buy Your Leased Car – Run These Two Numbers First

Should You Buy Your Leased Car – Run These Two Numbers First

September 11, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!