• bitcoinBitcoin(BTC)$76,741.00-0.76%
  • ethereumEthereum(ETH)$2,479.26-2.22%
  • tetherTether(USDT)$1.00-0.01%
  • binancecoinBNB(BNB)$716.07-2.70%
  • rippleXRP(XRP)$1.34-2.17%
  • usd-coinUSDC(USDC)$1.00-0.01%
  • solanaSolana(SOL)$99.72-2.25%
  • tronTRON(TRX)$0.3409780.04%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.00-1.59%
  • zcashZcash(ZEC)$1,095.61-4.71%
  • HyperliquidHyperliquid(HYPE)$77.40-3.61%
  • dogecoinDogecoin(DOGE)$0.083500-1.85%
  • RainRain(RAIN)$0.0153381.49%
  • moneroMonero(XMR)$538.110.88%
  • USDSUSDS(USDS)$1.00-0.01%
  • whitebitWhiteBIT Coin(WBT)$79.59-0.99%
  • chainlinkChainlink(LINK)$11.29-2.28%
  • leo-tokenLEO Token(LEO)$9.06-0.60%
  • cardanoCardano(ADA)$0.204574-2.23%
  • stellarStellar(XLM)$0.178346-2.15%
  • Ethena USDeEthena USDe(USDE)$1.00-0.02%
  • daiDai(DAI)$1.000.01%
  • bitcoin-cashBitcoin Cash(BCH)$223.38-3.29%
  • USD1USD1(USD1)$1.00-0.02%
  • litecoinLitecoin(LTC)$53.78-0.51%
  • uniswapUniswap(UNI)$6.27-1.50%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.35-1.89%
  • CantonCanton(CC)$0.094908-4.09%
  • Global DollarGlobal Dollar(USDG)$1.00-0.01%
  • hedera-hashgraphHedera(HBAR)$0.0752190.72%
  • avalanche-2Avalanche(AVAX)$7.32-1.69%
  • shiba-inuShiba Inu(SHIB)$0.000005-3.20%
  • nearNEAR Protocol(NEAR)$2.30-2.81%
  • suiSui(SUI)$0.71-2.52%
  • crypto-com-chainCronos(CRO)$0.0588291.36%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.01%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • tether-goldTether Gold(XAUT)$4,346.77-0.04%
  • MemeCoreMemeCore(M)$1.15-2.38%
  • Circle USYCCircle USYC(USYC)$1.140.00%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • okbOKB(OKB)$113.49-0.68%
  • BittensorBittensor(TAO)$232.50-1.70%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.14-0.06%
  • aaveAave(AAVE)$124.15-2.92%
  • pax-goldPAX Gold(PAXG)$4,351.93-0.04%
  • AsterAster(ASTER)$0.690.70%
  • mantleMantle(MNT)$0.56-2.86%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.0576301.03%
  • BitwayBitway(BTW)$0.6719.29%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

This AI Paper Demonstrates How Decoder-Only Transformers Mimic Infinite Multi-State Recurrent Neural Networks RNNs and Introduces TOVA for Enhanced Efficiency

January 15, 2024
in AI & Technology
Reading Time: 5 mins read
A A
This AI Paper Demonstrates How Decoder-Only Transformers Mimic Infinite Multi-State Recurrent Neural Networks RNNs and Introduces TOVA for Enhanced Efficiency
ShareShareShareShareShare

Transformers have taken over from recurrent neural networks (RNNs) as the preferred architecture for natural language processing (NLP). Transformers stand out conceptually because they directly access each token in a sequence, unlike RNNs that rely on maintaining a recurring state of past inputs. Decoders have emerged as a prominent variant within the realm of transformers. These decoders commonly produce output in an auto-regressive manner, meaning the generation of each token is influenced by the key and value computations of preceding tokens.

Researchers from The Hebrew University of Jerusalem and FAIR, AI at Meta, have demonstrated that the auto-regressive nature of transformers aligns with the fundamental principle of RNNs, which involves preserving a state from one step to the next. They formally redefine decoder-only transformers as multi-state RNNs (MSRNN), presenting a generalized version of traditional RNNs. This redefinition highlights that as the number of previous tokens increases during decoding, transformers become MSRNNs with infinite states. The researchers further show that transformers can be compressed into finite MSRNNs by limiting the number of tokens processed at each step. They introduce TOVA, a compression policy for MSRNNs, which selects tokens to retain based solely on their attention scores. The evaluation of TOVA is conducted on four long-range tasks.

https://arxiv.org/abs/2401.06104

The study compares transformers and RNNs, demonstrating that decoder-only transformers can be conceptualized as infinite multi-state RNNs, and pretrained transformers can be converted into finite multi-state RNNs by fixing the size of their hidden state. It reports perplexity on the PG-19 test set for language modeling. It uses test sets from the ZeroSCROLLS benchmark for evaluating long-range understanding, including long-range summarization and long-range question-answering tasks. The study mentions using the QASPER dataset for long text question answering and evaluating generated stories using GPT-4 as an evaluator.

https://arxiv.org/abs/2401.06104

The study demonstrates that decoder-only transformers can be conceptualized as infinite multi-state RNNs, and pretrained transformers can be converted into finite multi-state RNNs by fixing the size of their hidden state. The study also mentions modifying the attention mask to incorporate different MSRNN policies, such as the First In First Out (FIFO) strategy, to effectively parallel the language modeling task. The researchers use the GPT-4 model to evaluate the generated texts and compare the output of the TOVA policy with the topline model.

https://arxiv.org/abs/2401.06104

The study demonstrates that transformer decoder LLMs behave as finite MSRNNs even though they are trained as infinite MSRNNs. The proposed TOVA policy performs consistently better than other policies in long-range tasks with smaller cache sizes across all multi-state sizes and models. The experiments show that using TOVA with a quarter or even one-eighth of the full context yields results within one point of the topline model in language modeling tasks. The study also reports a significant reduction in LLM cache size, up to 88%, leading to reduced memory consumption during inference. The researchers acknowledge the computational constraints and approximate the infinite MSRNN with a sequence length of 4,096 tokens for extrapolation experiments.

To summarize, the researchers have redefined decoder transformers as multi-state RNNs with an infinite multi-state size. When the number of token representations that transformers can handle at each step is limited, it is the same as compressing it from infinite to finite MSRNNs. The TOVA policy, which is a simple compression method that selects which tokens to keep using their attention scores, has been found to outperform existing compression policies and performs comparably to the infinite MSRNN model with a reduced multi-state size. Although not trained, transformers often function as finite MSRNNs in practice. These findings provide insights into the inter-working of transformers and their connections to RNNs. Also, they have practical value in reducing the LLM cache size by up to 88%.


Check out the Paper. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter. Join our 36k+ ML SubReddit, 41k+ Facebook Community, Discord Channel, and LinkedIn Group.

If you like our work, you will love our newsletter..

Don’t Forget to join our Telegram Channel


YOU MAY ALSO LIKE

AWS Introduces Pizza Bot: An Open Source Inbox for Background AI Agents

Implementation of Machine Learning Workflows with NVIDIA cuML, RAPIDS, GPU Benchmarking, Explainability, Clustering, and Model Inference

Sana Hassan, a consulting intern at Marktechpost and dual-degree student at IIT Madras, is passionate about applying technology and AI to address real-world challenges. With a keen interest in solving practical problems, he brings a fresh perspective to the intersection of AI and real-life solutions.


[Free AI Event] 🐝 ‘Real-Time AI with Kafka and Streaming Data Analytics’ (Jan 15 2024, 10 am PST)


Credit: Source link

ShareTweetSendSharePin

Related Posts

AWS Introduces Pizza Bot: An Open Source Inbox for Background AI Agents
AI & Technology

AWS Introduces Pizza Bot: An Open Source Inbox for Background AI Agents

September 13, 2026
Implementation of Machine Learning Workflows with NVIDIA cuML, RAPIDS, GPU Benchmarking, Explainability, Clustering, and Model Inference
AI & Technology

Implementation of Machine Learning Workflows with NVIDIA cuML, RAPIDS, GPU Benchmarking, Explainability, Clustering, and Model Inference

September 13, 2026
Why Do Routers Have So Many Antennas?
AI & Technology

Why Do Routers Have So Many Antennas?

September 13, 2026
Hyundai Motor Group Puts Data Flywheel Into Full Operation – Unite.AI
AI & Technology

Hyundai Motor Group Puts Data Flywheel Into Full Operation – Unite.AI

September 13, 2026
Next Post
Teen speaks out after surviving Grand Canyon fall

Teen speaks out after surviving Grand Canyon fall

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Anthropic’s  Billion Credit Line Sets Stage for IPO

Anthropic’s $15 Billion Credit Line Sets Stage for IPO

September 8, 2026
Rep. Aguilar and James Blair weigh Iran war’s midterms impact

Rep. Aguilar and James Blair weigh Iran war’s midterms impact

September 13, 2026
Hunter Biden’s meme coin crushed 4 out of 5 buyers while one mystery trader raked in M

Hunter Biden’s meme coin crushed 4 out of 5 buyers while one mystery trader raked in $1M

September 10, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!