• bitcoinBitcoin(BTC)$78,620.00-0.94%
  • ethereumEthereum(ETH)$2,491.310.03%
  • tetherTether(USDT)$1.00-0.02%
  • binancecoinBNB(BNB)$755.911.55%
  • rippleXRP(XRP)$1.40-0.27%
  • usd-coinUSDC(USDC)$1.00-0.02%
  • solanaSolana(SOL)$103.55-1.27%
  • tronTRON(TRX)$0.3384110.53%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.060.00%
  • zcashZcash(ZEC)$1,137.85-4.34%
  • HyperliquidHyperliquid(HYPE)$84.33-3.34%
  • dogecoinDogecoin(DOGE)$0.0899090.39%
  • RainRain(RAIN)$0.0169982.80%
  • USDSUSDS(USDS)$1.00-0.01%
  • moneroMonero(XMR)$519.32-3.52%
  • chainlinkChainlink(LINK)$12.72-4.98%
  • whitebitWhiteBIT Coin(WBT)$78.597.49%
  • leo-tokenLEO Token(LEO)$9.21-0.57%
  • cardanoCardano(ADA)$0.2194540.54%
  • stellarStellar(XLM)$0.1912790.37%
  • bitcoin-cashBitcoin Cash(BCH)$257.140.35%
  • daiDai(DAI)$1.000.01%
  • uniswapUniswap(UNI)$7.121.94%
  • Ethena USDeEthena USDe(USDE)$1.000.00%
  • litecoinLitecoin(LTC)$55.67-0.27%
  • USD1USD1(USD1)$1.00-0.02%
  • CantonCanton(CC)$0.105097-2.79%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.40-1.56%
  • hedera-hashgraphHedera(HBAR)$0.080490-0.15%
  • avalanche-2Avalanche(AVAX)$8.081.85%
  • suiSui(SUI)$0.832.10%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • shiba-inuShiba Inu(SHIB)$0.000005-0.10%
  • nearNEAR Protocol(NEAR)$2.32-1.07%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.02%
  • crypto-com-chainCronos(CRO)$0.0586801.83%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • tether-goldTether Gold(XAUT)$4,392.42-0.21%
  • MemeCoreMemeCore(M)$1.162.74%
  • Circle USYCCircle USYC(USYC)$1.140.00%
  • BittensorBittensor(TAO)$254.65-5.54%
  • okbOKB(OKB)$116.352.70%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.14-0.01%
  • AsterAster(ASTER)$0.77-2.74%
  • mantleMantle(MNT)$0.62-2.22%
  • aaveAave(AAVE)$131.30-1.90%
  • pax-goldPAX Gold(PAXG)$4,396.44-0.23%
  • OndoOndo(ONDO)$0.381299-2.33%
  • polkadotPolkadot(DOT)$1.089.92%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Meet vLLM: An Open-Source Machine Learning Library for Fast LLM Inference and Serving

September 17, 2023
in AI & Technology
Reading Time: 4 mins read
A A
Meet vLLM: An Open-Source Machine Learning Library for Fast LLM Inference and Serving
ShareShareShareShareShare

Large language models (LLMs) have an ever-greater impact on how daily lives and careers are changing because they make possible new applications like programming assistants and universal chatbots. However, the operation of these applications comes at a substantial cost due to the significant hardware accelerator requirements, such as GPUs. Recent studies show that handling an LLM request can be expensive, up to ten times higher than a traditional keyword search. So, there is a growing need to boost the throughput of LLM serving systems to minimize the per-request expenses.

Performing high throughput serving of large language models (LLMs) requires batching sufficiently many requests at a time and the existing systems. 

However, existing systems need help because the key-value cache (KV cache) memory for each request is huge and can grow and shrink dynamically. It needs to be managed carefully, or when managed inefficiently, fragmentation and redundant duplication can greatly save this RAM, reducing the batch size.

The researchers have suggested PagedAttention, an attention algorithm inspired by the traditional virtual memory and paging techniques in operating systems, as a solution to this problem. To further reduce memory utilization, the researchers have also deployed vLLM. This LLM serving system provides almost zero waste in KV cache memory and flexible sharing of KV cache within and between requests.

vLLM utilizes PagedAttention to manage attention keys and values. By delivering up to 24 times more throughput than HuggingFace Transformers without requiring any changes to the model architecture, vLLM equipped with PagedAttention redefines the current state of the art in LLM serving.

Unlike conventional attention algorithms, they permit continuous key and value storage in non-contiguous memory space. PagedAttention divides each sequence’s KV cache into blocks, each with the keys and values for a predetermined amount of tokens. These blocks are efficiently identified by the PagedAttention kernel during the attention computation. As the blocks do not necessarily need to be contiguous, the keys and values can be managed flexibly.  

Memory leakage happens only in the ultimate block of a sequence within PagedAttention. In practical usage, this leads to effective memory utilization, with just a minimal 4% inefficiency. This enhancement in memory efficiency enables greater GPU utilization.

Also, PagedAttention has another key advantage of efficient memory sharing. PageAttention’s memory-sharing function considerably decreases the additional memory required for sampling techniques like parallel sampling and beam search. It can result in a speed gain of up to 2.2 times while reducing their memory utilization by up to 55%. This enhancement makes these sample techniques useful and effective for Large Language Model (LLM) services.

The researchers studied the accuracy of this system. They found that with the same amount of delay as cutting-edge systems like FasterTransformer and Orca, vLLM increases the throughput of well-known LLMs by 2-4. Larger models, more intricate decoding algorithms, and longer sequences result in a more noticeable improvement.


Check out the Paper, Github, and Reference Article. All Credit For This Research Goes To the Researchers on This Project. Also, don’t forget to join our 30k+ ML SubReddit, 40k+ Facebook Community, Discord Channel, and Email Newsletter, where we share the latest AI research news, cool AI projects, and more.

If you like our work, you will love our newsletter..


YOU MAY ALSO LIKE

Renault Is Building Its €17,900 Dacia Spring EV In Europe To Qualify For Local Subsidies

An Attractive ‘Mid-Size’ Foldable With Powerful Specs

Rachit Ranjan is a consulting intern at MarktechPost . He is currently pursuing his B.Tech from Indian Institute of Technology(IIT) Patna . He is actively shaping his career in the field of Artificial Intelligence and Data Science and is passionate and dedicated for exploring these fields.


🚀 The end of project management by humans (Sponsored)

Credit: Source link

ShareTweetSendSharePin

Related Posts

Renault Is Building Its €17,900 Dacia Spring EV In Europe To Qualify For Local Subsidies
AI & Technology

Renault Is Building Its €17,900 Dacia Spring EV In Europe To Qualify For Local Subsidies

September 8, 2026
An Attractive ‘Mid-Size’ Foldable With Powerful Specs
AI & Technology

An Attractive ‘Mid-Size’ Foldable With Powerful Specs

September 8, 2026
How Long Before a Real Crackdown on AI Model Decensoring? – Unite.AI
AI & Technology

How Long Before a Real Crackdown on AI Model Decensoring? – Unite.AI

September 8, 2026
Reducto Releases r-1: A Single Pass Document Parsing Model That Cuts Errors 20% at 1 Cent Per Page
AI & Technology

Reducto Releases r-1: A Single Pass Document Parsing Model That Cuts Errors 20% at 1 Cent Per Page

September 8, 2026
Next Post
Why Greycroft Is Investing in Optimus Ride

Why Greycroft Is Investing in Optimus Ride

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Mobile Games Designed to Be Addictive Get More Kid-Friendly

Mobile Games Designed to Be Addictive Get More Kid-Friendly

September 4, 2026
Nvidia Buys Hugging Face in .9 Billion Deal – The New York Times

Nvidia Buys Hugging Face in $12.9 Billion Deal – The New York Times

September 3, 2026
L.A. ‘graffiti towers’ to be cleaned in 90 days

L.A. ‘graffiti towers’ to be cleaned in 90 days

September 6, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!