• bitcoinBitcoin(BTC)$78,577.00-0.44%
  • ethereumEthereum(ETH)$2,488.040.11%
  • tetherTether(USDT)$1.00-0.01%
  • binancecoinBNB(BNB)$749.040.91%
  • rippleXRP(XRP)$1.411.30%
  • usd-coinUSDC(USDC)$1.000.01%
  • solanaSolana(SOL)$103.14-0.12%
  • tronTRON(TRX)$0.3390521.21%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.040.00%
  • zcashZcash(ZEC)$1,178.944.17%
  • HyperliquidHyperliquid(HYPE)$85.151.36%
  • dogecoinDogecoin(DOGE)$0.089547-0.57%
  • RainRain(RAIN)$0.015978-1.93%
  • USDSUSDS(USDS)$1.000.01%
  • whitebitWhiteBIT Coin(WBT)$81.386.52%
  • moneroMonero(XMR)$503.37-0.91%
  • chainlinkChainlink(LINK)$12.39-1.87%
  • leo-tokenLEO Token(LEO)$9.18-0.13%
  • cardanoCardano(ADA)$0.216359-0.96%
  • stellarStellar(XLM)$0.186827-1.43%
  • bitcoin-cashBitcoin Cash(BCH)$256.96-0.40%
  • daiDai(DAI)$1.000.00%
  • Ethena USDeEthena USDe(USDE)$1.000.01%
  • USD1USD1(USD1)$1.00-0.01%
  • CantonCanton(CC)$0.1079770.45%
  • uniswapUniswap(UNI)$6.80-3.30%
  • litecoinLitecoin(LTC)$53.93-2.56%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.38-0.07%
  • hedera-hashgraphHedera(HBAR)$0.078524-3.95%
  • avalanche-2Avalanche(AVAX)$7.94-1.48%
  • Global DollarGlobal Dollar(USDG)$1.00-0.01%
  • suiSui(SUI)$0.81-2.33%
  • shiba-inuShiba Inu(SHIB)$0.000005-1.36%
  • nearNEAR Protocol(NEAR)$2.27-0.72%
  • crypto-com-chainCronos(CRO)$0.0598295.23%
  • paypal-usdPayPal USD(PYUSD)$1.000.02%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • MemeCoreMemeCore(M)$1.202.08%
  • tether-goldTether Gold(XAUT)$4,374.40-1.20%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • BittensorBittensor(TAO)$255.69-0.46%
  • Ripple USDRipple USD(RLUSD)$1.000.01%
  • okbOKB(OKB)$113.82-2.39%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.140.40%
  • mantleMantle(MNT)$0.631.48%
  • AsterAster(ASTER)$0.75-1.82%
  • polkadotPolkadot(DOT)$1.1910.59%
  • aaveAave(AAVE)$128.15-2.34%
  • pax-goldPAX Gold(PAXG)$4,378.19-1.21%
  • Pump.funPump.fun(PUMP)$0.0044392.03%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

This AI Paper Proposes Retentive Networks (RetNet) as a Foundation Architecture for Large Language Models: Achieving Training Parallelism, Low-Cost Inference, and Good Performance

July 21, 2023
in AI & Technology
Reading Time: 4 mins read
A A
This AI Paper Proposes Retentive Networks (RetNet) as a Foundation Architecture for Large Language Models: Achieving Training Parallelism, Low-Cost Inference, and Good Performance
ShareShareShareShareShare

Transformer, which was first developed to address the sequential training problem with recurrent models, has since come to be accepted as the de facto architecture for big language models. Transformers’ O(N) complexity per step and memory-bound key-value cache make it unsuitable for deployment, trade-off training parallelism for poor inference. The sequence’s lengthening slows inference speed, increases latency, and uses more GPU memory. The next-generation architecture has continued extensive development to maintain training parallelism and competitive performance as Transformers while having effective O(1) inference. 

Figure 1: RetNet enables the “impossible triangle” to be achieved, which simultaneously achieves training parallelism, high performance, and cheap inference cost.

The so-called “impossible triangle” in Figure 1 illustrates how difficult it is to accomplish the objectives mentioned above simultaneously. Three key research streams have been present. To rewrite autoregressive inference in a recurrent form, linearized attention first approximates conventional attention scores exp(q . k) using kernels ϕ(q). ϕ(k). The method’s popularity could be improved because it performs and models less well than Transformers. The second strand forgoes parallel training in favor of recurrent models for effective inference. Element-wise operators are employed to fix acceleration, although this compromises representation capacity and performance. For attention, the third line of inquiry investigates substituting alternative mechanisms, such as S4 and its variations. 

There is no apparent winner compared to Transformers since none of the earlier works can escape the impasse. Researchers from  Microsoft Research and  Tsinghua University propose retentive networks (RetNet) which concurrently provide low-cost inference, effective long-sequence modeling, Transformer-comparable performance, and parallel model training. They specifically offer a multi-scale retention mechanism with three processing paradigms, similar, recurrent, and chunkwise recurrent representations, to replace multi-head attention. First, training parallelism may fully utilize GPU devices thanks to the parallel representation. Second, the recurrent representation makes efficient O(1) inference in terms of memory and computation possible. Both the deployment expense and latency may be greatly decreased. 

🚀 Build high-quality training datasets with Kili Technology and solve NLP machine learning challenges to develop powerful ML applications

Without key-value cache techniques, the method is also far more straightforward. Third, effective long-sequence modeling may be done using the chunkwise recurrent representation. They repeatedly encode the global blocks to conserve GPU memory while simultaneously encoding each local block to speed up processing. To compare RetNet with Transformer and its derivatives, they do comprehensive trials. According to experimental results on language modeling, RetNet constantly competes in terms of scaling curves and in-context learning. Additionally, RetNet’s inference cost is length-invariant. 

RetNet decodes 8.4 times quicker and uses 70% less memory than Transformers with key-value caches for a 7B model and an 8k sequence length. RetNet also saves 25–50% more memory while training accelerates compared to a normal Transformer and performs better than highly optimized FlashAttention. RetNet’s inference latency is unaffected by the batch size, enabling extremely high throughput. RetNet is a strong Transformer replacement for big language models because of its fascinating features.


Check out the Paper and Github link. All Credit For This Research Goes To the Researchers on This Project. Also, don’t forget to join our 26k+ ML SubReddit, Discord Channel, and Email Newsletter, where we share the latest AI research news, cool AI projects, and more.


YOU MAY ALSO LIKE

NSA, CISA, FBI Warn China-Based AI Firms Distill US Frontier Models – Unite.AI

How To Change And Customize Your Apple CarPlay Display

Aneesh Tickoo is a consulting intern at MarktechPost. He is currently pursuing his undergraduate degree in Data Science and Artificial Intelligence from the Indian Institute of Technology(IIT), Bhilai. He spends most of his time working on projects aimed at harnessing the power of machine learning. His research interest is image processing and is passionate about building solutions around it. He loves to connect with people and collaborate on interesting projects.


🔥 Gain a competitive
edge with data: Actionable market intelligence for global brands, retailers, analysts, and investors. (Sponsored)

Credit: Source link

ShareTweetSendSharePin

Related Posts

NSA, CISA, FBI Warn China-Based AI Firms Distill US Frontier Models – Unite.AI
AI & Technology

NSA, CISA, FBI Warn China-Based AI Firms Distill US Frontier Models – Unite.AI

September 9, 2026
How To Change And Customize Your Apple CarPlay Display
AI & Technology

How To Change And Customize Your Apple CarPlay Display

September 8, 2026
Is There Any Benefit To Restarting Your PC Regularly?
AI & Technology

Is There Any Benefit To Restarting Your PC Regularly?

September 8, 2026
Sierra Open-Sources Hyper-τ-Bench, a Benchmark for Agent Construction
AI & Technology

Sierra Open-Sources Hyper-τ-Bench, a Benchmark for Agent Construction

September 8, 2026
Next Post
U.S. Stocks Will Likely Struggle Till Second Half Says Ameriprise Strategist

U.S. Stocks Will Likely Struggle Till Second Half Says Ameriprise Strategist

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
My Sister Thinks I’m Out To Kill Her (I Just Want To Sell a House)

My Sister Thinks I’m Out To Kill Her (I Just Want To Sell a House)

September 2, 2026
You Make 285,000 And Have Nothing To Show For It

You Make 285,000 And Have Nothing To Show For It

September 6, 2026
Buy 3 Ideal September Dividend Dogs Out Of Barron’s 58 August Picks

Buy 3 Ideal September Dividend Dogs Out Of Barron’s 58 August Picks

September 5, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!