• bitcoinBitcoin(BTC)$77,282.00-0.73%
  • ethereumEthereum(ETH)$2,539.160.92%
  • tetherTether(USDT)$1.000.02%
  • binancecoinBNB(BNB)$736.002.07%
  • rippleXRP(XRP)$1.37-0.10%
  • usd-coinUSDC(USDC)$1.000.00%
  • solanaSolana(SOL)$101.890.15%
  • tronTRON(TRX)$0.3403941.28%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.02-1.77%
  • zcashZcash(ZEC)$1,154.12-0.32%
  • HyperliquidHyperliquid(HYPE)$80.28-1.67%
  • dogecoinDogecoin(DOGE)$0.084952-0.18%
  • RainRain(RAIN)$0.015074-5.01%
  • moneroMonero(XMR)$531.144.03%
  • USDSUSDS(USDS)$1.000.00%
  • whitebitWhiteBIT Coin(WBT)$80.37-0.58%
  • chainlinkChainlink(LINK)$11.55-1.43%
  • leo-tokenLEO Token(LEO)$9.12-0.35%
  • cardanoCardano(ADA)$0.2084280.24%
  • stellarStellar(XLM)$0.1811791.35%
  • bitcoin-cashBitcoin Cash(BCH)$230.480.56%
  • Ethena USDeEthena USDe(USDE)$1.000.03%
  • daiDai(DAI)$1.00-0.02%
  • USD1USD1(USD1)$1.000.01%
  • litecoinLitecoin(LTC)$53.880.61%
  • uniswapUniswap(UNI)$6.413.38%
  • CantonCanton(CC)$0.097841-0.10%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.370.57%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • avalanche-2Avalanche(AVAX)$7.44-1.32%
  • hedera-hashgraphHedera(HBAR)$0.074580-1.09%
  • shiba-inuShiba Inu(SHIB)$0.0000052.62%
  • nearNEAR Protocol(NEAR)$2.37-9.57%
  • suiSui(SUI)$0.73-1.13%
  • crypto-com-chainCronos(CRO)$0.0582662.78%
  • paypal-usdPayPal USD(PYUSD)$1.000.00%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • MemeCoreMemeCore(M)$1.17-1.63%
  • tether-goldTether Gold(XAUT)$4,347.99-0.93%
  • Circle USYCCircle USYC(USYC)$1.140.00%
  • Ripple USDRipple USD(RLUSD)$1.00-0.01%
  • okbOKB(OKB)$113.920.73%
  • BittensorBittensor(TAO)$234.80-1.84%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.150.04%
  • aaveAave(AAVE)$126.281.04%
  • mantleMantle(MNT)$0.57-2.66%
  • pax-goldPAX Gold(PAXG)$4,353.66-0.91%
  • AsterAster(ASTER)$0.68-1.75%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.0569418.00%
  • polkadotPolkadot(DOT)$1.04-3.56%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Google AI Research Introduces GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints

January 31, 2024
in AI & Technology
Reading Time: 4 mins read
A A
Google AI Research Introduces GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints
ShareShareShareShareShare

In the enchanting world of language models and attention mechanisms, picture a daring quest to accelerate decoder inference and enhance the prowess of large language models. Our tale unfolds with the discovery of multi-query attention (MQA), a captivating technique that promises speedier results. Multi-query attention (MQA) expedites decoder inference through the employment of a single key-value head. 

However, its efficiency is countered by the potential for a decline in quality. Furthermore, there may be hesitation in training a separate model solely dedicated to hastening inference. Despite its benefits, the use of MQA is linked with drawbacks such as quality degradation and training instability. Moreover, the feasibility of developing distinct models optimized for both quality and inference is questioned due to potential limitations.

The above figure demonstrates the overview of conversion from multi-head to multi-query attention. Key and value projection matrices from all heads are mean pooled into a single head. 

The paper introduces two contributions aimed at enhancing the efficiency of large language models during inference. Firstly, it demonstrates that language model checkpoints employing multi-head attention (MHA) can be uptrained, as outlined by Komatsuzaki et al. in 2022, to incorporate multi-query attention (MQA) with a minimal fraction of the original training compute. This approach offers a cost-effective means of obtaining both rapid multi-query functionality and high-quality MHA checkpoints.

Secondly, the paper suggests grouped-query attention (GQA) as an interpolation between multi-head and multi-query attention, utilizing single key and value heads for each subgroup of query heads. The research illustrates that uptrained GQA achieves quality levels close to multi-head attention while maintaining a speed comparable to that of multi-query attention.

Employing language models for swift responses becomes expensive due to the high memory demand for loading keys and values. Although multi-query attention addresses this issue by cutting down on memory usage, it does so at the cost of reducing the model size and accuracy. The proposed approach involves transforming multi-head attention models into multi-query models using only a fraction of the original training. Furthermore, the introduction of grouped-query attention, a combination of multi-query and multi-head attention, maintains quality comparable to multi-head attention while operating at a speed nearly as fast as multi-query attention.

In conclusion, the objective of this paper is to enhance the efficiency of language models in handling substantial amounts of information while minimising computer memory usage. This is particularly crucial when dealing with longer sequences, where assessing quality poses challenges. The evaluation for summarization involves using a metric called Rouge score, with an acknowledgment of its imperfect nature. Due to certain limitations in the testing methodology, the certainty of the correctness of our choices is not absolute.

Additionally, a direct comparison of our XXL GQA model with a counterpart trained from scratch was not conducted, preventing a clear understanding of its performance relative to starting anew. Lastly, the evaluations focused exclusively on models engaged in both reading and generating information. There are other popular models dedicated solely to information generation, and there is a belief that our GQA approach may prove more effective for them compared to an alternative technique known as MQA.


Check out the Paper. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter. Join our 36k+ ML SubReddit, 41k+ Facebook Community, Discord Channel, and LinkedIn Group.

If you like our work, you will love our newsletter..

Don’t Forget to join our Telegram Channel


YOU MAY ALSO LIKE

Is A 256GB SSD Better Than A 1TB Hard Drive? It Depends How You’re Using It

What Is Benchmark Saturation? Why Yesterday’s AI Tests Stop Working – Unite.AI

Janhavi Lande, is an Engineering Physics graduate from IIT Guwahati, class of 2023. She is an upcoming data scientist and has been working in the world of ml/ai research for the past two years. She is most fascinated by this ever changing world and its constant demand of humans to keep up with it. In her pastime she enjoys traveling, reading and writing poems.


🎯 [FREE AI WEBINAR] ‘Create Embeddings on Real-Time Data with OpenAI & SingleStore Job Service’ (Jan 31, 2024)


Credit: Source link

ShareTweetSendSharePin

Related Posts

Is A 256GB SSD Better Than A 1TB Hard Drive? It Depends How You’re Using It
AI & Technology

Is A 256GB SSD Better Than A 1TB Hard Drive? It Depends How You’re Using It

September 12, 2026
What Is Benchmark Saturation? Why Yesterday’s AI Tests Stop Working – Unite.AI
AI & Technology

What Is Benchmark Saturation? Why Yesterday’s AI Tests Stop Working – Unite.AI

September 12, 2026
Kai-Fu Lee Says China Will Win AI Reach Race
AI & Technology

Kai-Fu Lee Says China Will Win AI Reach Race

September 12, 2026
Everybody’s Business: Unpacking Apple’s Upcoming Launches
AI & Technology

Everybody’s Business: Unpacking Apple’s Upcoming Launches

September 12, 2026
Next Post
Trump continues on the campaign trail amid the threat of another indictment

Trump continues on the campaign trail amid the threat of another indictment

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
How These XL Phones Compete

How These XL Phones Compete

September 10, 2026
Anthropic Caught Scientists Using Claude To Further Biological Weapon Research

Anthropic Caught Scientists Using Claude To Further Biological Weapon Research

September 10, 2026
OpenAI Rolls Out Its Most Advanced Model Yet

OpenAI Rolls Out Its Most Advanced Model Yet

September 7, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!