• bitcoinBitcoin(BTC)$79,242.001.16%
  • ethereumEthereum(ETH)$2,509.701.21%
  • tetherTether(USDT)$1.000.02%
  • binancecoinBNB(BNB)$745.70-0.51%
  • rippleXRP(XRP)$1.431.34%
  • usd-coinUSDC(USDC)$1.000.01%
  • solanaSolana(SOL)$103.960.86%
  • tronTRON(TRX)$0.338523-0.24%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.043.22%
  • zcashZcash(ZEC)$1,291.919.90%
  • HyperliquidHyperliquid(HYPE)$86.503.49%
  • dogecoinDogecoin(DOGE)$0.0903090.97%
  • RainRain(RAIN)$0.016447-1.31%
  • USDSUSDS(USDS)$1.000.03%
  • whitebitWhiteBIT Coin(WBT)$81.902.72%
  • moneroMonero(XMR)$511.463.15%
  • chainlinkChainlink(LINK)$12.12-3.21%
  • leo-tokenLEO Token(LEO)$9.19-0.29%
  • cardanoCardano(ADA)$0.219394-0.38%
  • stellarStellar(XLM)$0.188307-0.49%
  • bitcoin-cashBitcoin Cash(BCH)$259.000.99%
  • daiDai(DAI)$1.000.00%
  • Ethena USDeEthena USDe(USDE)$1.000.01%
  • USD1USD1(USD1)$1.00-0.01%
  • litecoinLitecoin(LTC)$54.33-0.58%
  • CantonCanton(CC)$0.1052790.64%
  • uniswapUniswap(UNI)$6.65-3.99%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.39-0.45%
  • hedera-hashgraphHedera(HBAR)$0.078807-1.54%
  • avalanche-2Avalanche(AVAX)$7.95-0.19%
  • nearNEAR Protocol(NEAR)$2.6212.92%
  • suiSui(SUI)$0.810.25%
  • Global DollarGlobal Dollar(USDG)$1.00-0.01%
  • shiba-inuShiba Inu(SHIB)$0.0000050.12%
  • crypto-com-chainCronos(CRO)$0.059832-1.17%
  • paypal-usdPayPal USD(PYUSD)$1.000.02%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • tether-goldTether Gold(XAUT)$4,420.730.67%
  • MemeCoreMemeCore(M)$1.180.78%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • BittensorBittensor(TAO)$263.433.69%
  • Ripple USDRipple USD(RLUSD)$1.000.01%
  • okbOKB(OKB)$114.01-0.34%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.150.03%
  • mantleMantle(MNT)$0.642.72%
  • AsterAster(ASTER)$0.75-1.09%
  • aaveAave(AAVE)$130.221.16%
  • Pump.funPump.fun(PUMP)$0.0046815.73%
  • polkadotPolkadot(DOT)$1.133.51%
  • pax-goldPAX Gold(PAXG)$4,424.910.67%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

How Can We Efficiently Deploy Large Language Models in Streaming Applications? This AI Paper Introduces the StreamingLLM Framework for Infinite Sequence Lengths

October 10, 2023
in AI & Technology
Reading Time: 5 mins read
A A
How Can We Efficiently Deploy Large Language Models in Streaming Applications? This AI Paper Introduces the StreamingLLM Framework for Infinite Sequence Lengths
ShareShareShareShareShare

Large Language Models (LLMs) are increasingly used to power natural language processing applications, including code completion, question answering, document summarization, and dialogue systems. Pretrained LLMs must be capable of performing extended sequence creation precisely and quickly to reach their full potential. An ideal ChatBot helper, for instance, can reliably edit the content of recent day-long chats. To generalize to greater sequence lengths than they have been pretrained on, such as 4K for Llama-2, is very difficult for LLM. Because of the attention window during pre-training, LLMs are restricted. 

Although significant attempts have been made to increase the size of this window and increase training and inference effectiveness for long inputs, the permissible sequence length still needs to be revised, which prevents permanent deployments. Researchers from MIT, Meta AI and Carnegie Mellon University initially discuss the idea of LLM streaming applications in this study and pose the following query: Two main issues emerge when using LLMs for endless input streams: 

1. Transformer-based LLMs cache the Key and Value states (KV) of all prior tokens during the decoding stage, as shown in Figure 1(a), which may result in excessive memory use and a rise in decoding delay. 

2. The performance of existing models suffers when the duration of the sequence exceeds the attention window size determined during pre-training. 

Figure 1 compares StreamingLLM to previous techniques. The Tth token (T >> L) is predicted by the language model, which has been pre-trained on texts of length L. (a) Dense Attention has a rising cache capacity and an O(T^2) time complexity. When the text length is more than the pre-training text length, its performance suffers. (b) Window Attention stores the KV of the newest L tokens in its cache. Although performance is good for inference, it rapidly deteriorates when the keys and values of the initial tokens are removed. For each new token, (c) Sliding Window with Re-computation reconstructs the KV states using the L most recent tokens. Although it excels at handling lengthy texts, due to its O(T L^2 ) complexity and quadratic attention in context re-computation, it is incredibly sluggish. (d) For steady attention computation, StreamingLLM retains the attention sink (a few beginning tokens), together with the most recent tokens. It works effectively and consistently with long texts. The Llama-2-13B model is used to calculate perplexities for the first book (65K tokens) in the PG-19 test set.

Window attention is an obvious strategy that keeps a fixed-size sliding window on the KV states of the most recent tokens (Figure 1b). Even merely evicting the KV of the first token causes the model to collapse after the sequence length exceeds the cache capacity, even if it guarantees consistent memory use and decoding performance after the cache is first full. A further tactic is a sliding window with recomputation (Figure 1c), which reconstructs the KV states of recent tokens for each created token. The calculation of quadratic attention within its window makes this technique much slower, even if it performs well, making it unsuitable for real-world streaming applications. 

They discover intriguing phenomena of autoregressive LLMs to explain the failure of window attention: a startlingly high attention score is allotted to the initial tokens, regardless of their relevance to the language modeling job. These tokens are referred to as “attention sinks.” They receive significant attention scores while having little semantic value. The Softmax operation, which demands that attention scores add up to one for all contextual tokens, is cited as the cause. As a result, the model must assign these extra attention values to add up to one, even when the current query does not have a good match in many earlier tokens. 

Initial tokens are used as attention sinks for a simple reason: they are visible to practically all subsequent tokens due to the nature of autoregressive language modeling, making them easier to train. They suggest StreamingLLM, a straightforward and effective architecture that enables LLMs prepared with a finite attention window to work on text of indefinite duration without fine-tuning, in light of the abovementioned discoveries. Because attention drains have high attention values, StreamingLLM uses this property to keep the attention score distribution reasonably regular. StreamingLLM maintains the KVs of the sliding window and the attention sink tokens (with only four initial tokens needed) to anchor the attention computation and stabilize the model’s performance. 

Models like Llama-2-B, MPT-B, Falcon-B, and PythiaB can accurately represent 4 million tokens with the help of StreamingLLM, and maybe much more. StreamingLLM achieves up to 22.2 speedups compared to the only practical baseline, sliding window with recomputation, realizing the streaming usage of LLMs. Finally, they show that language models may be pre-trained to require only a single attention sink token for streaming deployment, confirming their attention sink hypothesis. They propose that a selected attention sink can be implemented as an additional learnable token at the start of each training sample. Introducing this single sink token maintains the model’s performance in streaming instances by pre-training language models with 160 million parameters from scratch. This contrasts with vanilla models, which call for reintroducing several initial tokens as attention sinks to maintain the same degree of performance.


Check out the Paper. All Credit For This Research Goes To the Researchers on This Project. Also, don’t forget to join our 31k+ ML SubReddit, 40k+ Facebook Community, Discord Channel, and Email Newsletter, where we share the latest AI research news, cool AI projects, and more.

If you like our work, you will love our newsletter..

We are also on WhatsApp. Join our AI Channel on Whatsapp..


YOU MAY ALSO LIKE

Why It’s Time to Abandon the ‘Set It and Forget It’ Model – Unite.AI

Lyft Is Now Offering Waymo Rides In Nashville

Aneesh Tickoo is a consulting intern at MarktechPost. He is currently pursuing his undergraduate degree in Data Science and Artificial Intelligence from the Indian Institute of Technology(IIT), Bhilai. He spends most of his time working on projects aimed at harnessing the power of machine learning. His research interest is image processing and is passionate about building solutions around it. He loves to connect with people and collaborate on interesting projects.


▶️ Now Watch AI Research Updates On Our Youtube Channel [Watch Now]

Credit: Source link

ShareTweetSendSharePin

Related Posts

Why It’s Time to Abandon the ‘Set It and Forget It’ Model – Unite.AI
AI & Technology

Why It’s Time to Abandon the ‘Set It and Forget It’ Model – Unite.AI

September 9, 2026
Lyft Is Now Offering Waymo Rides In Nashville
AI & Technology

Lyft Is Now Offering Waymo Rides In Nashville

September 9, 2026
Harvey Secures 0M in Fresh Funding, Valuation Climbs to .5B – Unite.AI
AI & Technology

Harvey Secures $550M in Fresh Funding, Valuation Climbs to $15.5B – Unite.AI

September 9, 2026
How To Take Full Advantage Of Gemini When Planning Your Next Trip
AI & Technology

How To Take Full Advantage Of Gemini When Planning Your Next Trip

September 9, 2026
Next Post
Bloomberg Technology 08/10/2023

Bloomberg Technology 08/10/2023

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
LIVE: Trump makes announcement on nuclear innovation | NBC News

LIVE: Trump makes announcement on nuclear innovation | NBC News

September 5, 2026
Wayve, Uber Bring Robotaxi Rides to London

Wayve, Uber Bring Robotaxi Rides to London

September 3, 2026
How To Reset The Camera Settings On Your iPhone

How To Reset The Camera Settings On Your iPhone

September 8, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!