• bitcoinBitcoin(BTC)$77,033.00-0.32%
  • ethereumEthereum(ETH)$2,490.93-1.31%
  • tetherTether(USDT)$1.00-0.02%
  • binancecoinBNB(BNB)$719.11-1.21%
  • rippleXRP(XRP)$1.35-1.26%
  • usd-coinUSDC(USDC)$1.00-0.01%
  • solanaSolana(SOL)$99.99-2.16%
  • tronTRON(TRX)$0.338248-0.71%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.000.00%
  • zcashZcash(ZEC)$1,073.33-4.65%
  • HyperliquidHyperliquid(HYPE)$78.12-1.55%
  • dogecoinDogecoin(DOGE)$0.083073-2.17%
  • RainRain(RAIN)$0.015207-3.40%
  • USDSUSDS(USDS)$1.00-0.01%
  • moneroMonero(XMR)$517.85-4.49%
  • whitebitWhiteBIT Coin(WBT)$79.82-0.65%
  • chainlinkChainlink(LINK)$11.30-1.84%
  • leo-tokenLEO Token(LEO)$9.02-0.80%
  • cardanoCardano(ADA)$0.204910-1.45%
  • stellarStellar(XLM)$0.178753-0.85%
  • Ethena USDeEthena USDe(USDE)$1.00-0.03%
  • daiDai(DAI)$1.000.00%
  • bitcoin-cashBitcoin Cash(BCH)$222.66-1.57%
  • USD1USD1(USD1)$1.00-0.01%
  • litecoinLitecoin(LTC)$54.210.71%
  • uniswapUniswap(UNI)$6.23-2.22%
  • CantonCanton(CC)$0.096625-1.46%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.35-2.32%
  • hedera-hashgraphHedera(HBAR)$0.0757680.94%
  • Global DollarGlobal Dollar(USDG)$1.00-0.01%
  • avalanche-2Avalanche(AVAX)$7.36-0.50%
  • shiba-inuShiba Inu(SHIB)$0.000005-1.97%
  • nearNEAR Protocol(NEAR)$2.33-1.35%
  • suiSui(SUI)$0.71-2.81%
  • crypto-com-chainCronos(CRO)$0.057793-2.97%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.01%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • tether-goldTether Gold(XAUT)$4,351.340.01%
  • Circle USYCCircle USYC(USYC)$1.140.00%
  • MemeCoreMemeCore(M)$1.13-4.99%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • okbOKB(OKB)$113.87-0.54%
  • BittensorBittensor(TAO)$235.030.72%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.15-0.09%
  • aaveAave(AAVE)$125.12-0.98%
  • pax-goldPAX Gold(PAXG)$4,355.630.00%
  • AsterAster(ASTER)$0.690.56%
  • mantleMantle(MNT)$0.560.67%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.056953-0.17%
  • BitwayBitway(BTW)$0.6619.14%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Google DeepMind Introduces Tandem Transformers for Inference Efficient Large Language Models LLMs

March 2, 2024
in AI & Technology
Reading Time: 4 mins read
A A
Google DeepMind Introduces Tandem Transformers for Inference Efficient Large Language Models LLMs
ShareShareShareShareShare

Very large language models (LLMs) continue to face major computational cost barriers, which prevents their broad deployment, even with inference optimization approaches that have advanced significantly. Sequentially producing tokens throughout the autoregressive generation process is a major cause of the high inference latency. Because ML accelerators (GPUs/TPUs) are designed for matrix-matrix multiplications and not the matrix-vector operations common in LLMs, this limitation prevents them from being fully utilized. As a result, autoregressive answer creation is far less efficient than prompt processing, which involves handling all tokens concurrently. 

However, the relative importance of the ability to comprehend the query or prefill (natural language understanding, or NLU) and the ability to produce an answer (natural language generation, or NLG) remains unclear. Modern LLM designs that rely solely on decoders bind these two activities together.

A new study by Google Research and DeepMind takes an efficiency-oriented look at this basic question. Their study presents Tandem Transformers, a new design that gives NLU (prefill processing) a far larger share of the model’s resources than NLG (response generation) does.  

The researchers implement a projection layer to bring the perhaps higher-dimensional representation space into alignment. Experiments with Tandem (PaLM2-Bison, PaLM2-Gecko) show that the capacity required for NLU vs NLG parts of LLMs can be separated, resulting in a more efficient design without a noticeable decrease in accuracy (where PaLM2-Gecko < PaLM2-Otter < PaLM2-Bison, according to model size). To maintain high accuracy, Tandem’s primary model refreshes all prefill representations, in contrast to an encoder-decoder architecture that would process query/prefix through an encoder and then generate the entire response through a decoder. 

They recommend Tandem + SPEED for applications that want output indistinguishable from the main model. The speculative decoding (SPEED) framework uses the Tandem small model to create draft tokens. Then, the large model verifies them. Improving draft quality while decreasing verification overhead relative to traditional SPEED is greatly aided by Tandem’s small model’s capacity to respond to the representations of large models.

Since Tandem is an independent model, it can produce respectable results without inherently requiring verification by a huge model. Tandem + SPEED can also leverage ML representations while autoregressively generating tokens, giving the drafter a far better compromise between token quality and model latency. Studies have demonstrated that logit distillation is useful for improving SPEED draft model training. This method works well with distillation and is complementary to it. Empirical Results for Tandem + SPEED. Lastly, they evaluate TPUv5e’s latency extensively for both the stand-alone and SPEED Tandem versions (PaLM2- Bison, PaLM2-Gecko), where PaLM2- Bison is the main large model and PaLM2- Gecko is the secondary small model. The researchers find that Tandem + SPEED with distillation can outperform the baseline PaLM2-Bison model by a factor of at least 2.19 on various datasets while maintaining the same output quality. As a bonus, their model is 1.11 to 1.17 times faster than the usual SPEED with the small model as the secondary model. Using an adaptive block length in SPEED, Tandem’s latency can be further reduced on various datasets by 1.04× to 1.09×.


Check out the Paper. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter and Google News. Join our 38k+ ML SubReddit, 41k+ Facebook Community, Discord Channel, and LinkedIn Group.

If you like our work, you will love our newsletter..

Don’t Forget to join our Telegram Channel

You may also like our FREE AI Courses….


YOU MAY ALSO LIKE

Which Is Better For Charging Your MacBook?

Nadella Announces Public Consultation on Microsoft’s MAI Model Rules – Unite.AI

Dhanshree Shenwai is a Computer Science Engineer and has a good experience in FinTech companies covering Financial, Cards & Payments and Banking domain with keen interest in applications of AI. She is enthusiastic about exploring new technologies and advancements in today’s evolving world making everyone’s life easy.


🐝 Join the Fastest Growing AI Research Newsletter Read by Researchers from Google + NVIDIA + Meta + Stanford + MIT + Microsoft and many others…


Credit: Source link

ShareTweetSendSharePin

Related Posts

Which Is Better For Charging Your MacBook?
AI & Technology

Which Is Better For Charging Your MacBook?

September 14, 2026
Nadella Announces Public Consultation on Microsoft’s MAI Model Rules – Unite.AI
AI & Technology

Nadella Announces Public Consultation on Microsoft’s MAI Model Rules – Unite.AI

September 13, 2026
How To Fix iMessage “Not Delivered” Error On iPhones
AI & Technology

How To Fix iMessage “Not Delivered” Error On iPhones

September 13, 2026
How To Adjust The Liquid Glass Effect On Your iPhone With iOS 27
AI & Technology

How To Adjust The Liquid Glass Effect On Your iPhone With iOS 27

September 13, 2026
Next Post
Breaking down Jack Smith’s ‘textbook response’ to Trump indictment

Breaking down Jack Smith's 'textbook response' to Trump indictment

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Everything You Need to Know About Apple’s iPhone Duo

Everything You Need to Know About Apple’s iPhone Duo

September 12, 2026
Alamos Gold: The Mine That Broke Isn’t The One That Matters (NYSE:AGI)

Alamos Gold: The Mine That Broke Isn’t The One That Matters (NYSE:AGI)

September 9, 2026
Viral video ‘lunatic’ jumps in front of Tesla Cybercab, forcing brakes and raising safety questions

Viral video ‘lunatic’ jumps in front of Tesla Cybercab, forcing brakes and raising safety questions

September 8, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!