• bitcoinBitcoin(BTC)$79,542.00-1.90%
  • ethereumEthereum(ETH)$2,450.72-2.67%
  • tetherTether(USDT)$1.000.02%
  • binancecoinBNB(BNB)$721.90-0.62%
  • rippleXRP(XRP)$1.40-3.79%
  • usd-coinUSDC(USDC)$1.000.01%
  • solanaSolana(SOL)$101.89-2.06%
  • tronTRON(TRX)$0.3317790.75%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.040.53%
  • HyperliquidHyperliquid(HYPE)$83.93-3.56%
  • zcashZcash(ZEC)$1,019.117.24%
  • dogecoinDogecoin(DOGE)$0.084668-3.10%
  • RainRain(RAIN)$0.016418-4.15%
  • moneroMonero(XMR)$533.846.07%
  • USDSUSDS(USDS)$1.00-0.01%
  • chainlinkChainlink(LINK)$11.66-2.67%
  • whitebitWhiteBIT Coin(WBT)$73.11-1.25%
  • leo-tokenLEO Token(LEO)$9.22-1.00%
  • cardanoCardano(ADA)$0.210770-6.69%
  • stellarStellar(XLM)$0.180926-1.66%
  • bitcoin-cashBitcoin Cash(BCH)$248.30-3.30%
  • daiDai(DAI)$1.000.01%
  • Ethena USDeEthena USDe(USDE)$1.000.02%
  • CantonCanton(CC)$0.108147-3.54%
  • USD1USD1(USD1)$1.00-0.01%
  • litecoinLitecoin(LTC)$52.382.19%
  • uniswapUniswap(UNI)$6.30-0.19%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.401.70%
  • hedera-hashgraphHedera(HBAR)$0.0787840.33%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • avalanche-2Avalanche(AVAX)$7.41-1.32%
  • suiSui(SUI)$0.76-1.45%
  • shiba-inuShiba Inu(SHIB)$0.000005-1.71%
  • nearNEAR Protocol(NEAR)$2.2314.19%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.01%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • tether-goldTether Gold(XAUT)$4,426.19-0.88%
  • crypto-com-chainCronos(CRO)$0.055696-4.38%
  • Circle USYCCircle USYC(USYC)$1.140.04%
  • MemeCoreMemeCore(M)$1.139.02%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • okbOKB(OKB)$109.23-0.49%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.14-0.28%
  • BittensorBittensor(TAO)$227.92-0.93%
  • aaveAave(AAVE)$129.75-3.44%
  • AsterAster(ASTER)$0.731.77%
  • pax-goldPAX Gold(PAXG)$4,432.87-0.92%
  • mantleMantle(MNT)$0.570.35%
  • OndoOndo(ONDO)$0.3708981.56%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.056442-2.94%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Meet vLLM: An Open-Source LLM Inference And Serving Library That Accelerates HuggingFace Transformers By 24x

June 25, 2023
in AI & Technology
Reading Time: 4 mins read
A A
Meet vLLM: An Open-Source LLM Inference And Serving Library That Accelerates HuggingFace Transformers By 24x
ShareShareShareShareShare

Large language models, or LLMs in short, have emerged as a groundbreaking advancement in the field of artificial intelligence (AI). These models, such as GPT-3, have completely revolutionalized natural language understanding. With the capacity of such models to interpret vast amounts of existing data and generate human-like texts, these models hold immense potential to shape the future of AI and open up new possibilities for human-machine interaction and communication. However, despite the massive success achieved by LLMs, one significant challenge often associated with such models is their computational inefficiency, leading to slow performance even on the most powerful hardware. Since these models comprise millions and billions of parameters, training such models demands extensive computational resources, memory, and processing power, which is not always accessible. Moreover, these complex architectures with slow response times can make LLMs impractical for real-time or interactive applications. As a result, addressing these challenges becomes essential in unlocking the full potential of LLMs and making their benefits more widely accessible. 

Tacking this problem statement, researchers from the University of California, Berkeley, have developed vLLM, an open-source library that is a simpler, faster, and cheaper alternative for LLM inference and serving. Large Model Systems Organization (LMSYS) is currently using the library to power their Vicuna and Chatbot Arena. By switching to vLLM as their backend, in contrast to the initial HuggingFace Transformers based backend, the research organization has managed to handle peak traffic efficiently (5 times more than before) while using limited computational resources and reducing high operational costs. Currently, vLLM supports several HuggingFace models like GPT-2, GPT BigCode, and LLaMA, to name a few. It achieves throughput levels that are 24 times higher than those of HuggingFace Transformers while maintaining the same model architecture and without necessitating any modifications.

As a part of their preliminary research, the Berkeley researchers determined that memory-related issues pose the primary constraint on the performance of LLMs. LLMs use input tokens to generate attention key and value tensors, which are then cached in GPU memory for generating subsequent tokens. These dynamic key and value tensors, known as KV cache, occupy a substantial portion of memory, and managing them becomes a cumbersome task. To address this challenge, the researchers introduced the innovative concept of PagedAttention, a novel attention algorithm that extends the conventional idea of paging in operating systems to LLM serving. PagedAttention offers a more flexible approach to managing key and value tensors by storing them in non-contiguous memory spaces, eliminating the requirement for continuous long memory blocks. These blocks can be independently retrieved using a block table during attention computation, leading to more efficient memory utilization. Adopting this clever technique reduces memory wastage to less than 4%, resulting in near-optimal memory usage. Moreover, PagedAttention can batch 5x more sequences together, thereby enhancing GPU utilization and throughput.

🔥 Unleash the power of Live Proxies: Private, undetectable residential and mobile IPs.

PagedAttention offers the additional benefit of efficient memory sharing. During parallel sampling, i.e., when multiple output sequences are created simultaneously from a single prompt, PagedAttention enables the sharing of computational resources and memory associated with that prompt. This is accomplished by utilizing a block table, where different sequences within PagedAttention can share blocks by mapping logical blocks to the same physical block. By employing this memory-sharing mechanism, PagedAttention not only minimizes memory usage but also ensures secure sharing. The experimental evaluations conducted by the researchers revealed that parallel sampling could reduce memory usage by a whopping 55%, resulting in a 2.2 times increase in throughput.

To summarize, vLLM effectively handles the management of attention key and value memory through the implementation of the PagedAttention mechanism. This results in exceptional throughput performance. Moreover, vLLM seamlessly integrates with well-known HuggingFace models and can be utilized alongside different decoding algorithms, such as parallel sampling. The library can be installed using a simple pip command and is currently available for both offline inference and online serving.


Check Out The Blog Article and Github. Don’t forget to join our 25k+ ML SubReddit, Discord Channel, and Email Newsletter, where we share the latest AI research news, cool AI projects, and more. If you have any questions regarding the above article or if we missed anything, feel free to email us at [email protected]

🚀 Check Out 100’s AI Tools in AI Tools Club


YOU MAY ALSO LIKE

How To See What’s Taking Up Space On Your Windows PC

The Tetris Company Wants Nothing To Do With The White House’s New Copycat Game

Khushboo Gupta is a consulting intern at MarktechPost. She is currently pursuing her B.Tech from the Indian Institute of Technology(IIT), Goa. She is passionate about the fields of Machine Learning, Natural Language Processing and Web Development. She enjoys learning more about the technical field by participating in several challenges.


Credit: Source link

ShareTweetSendSharePin

Related Posts

How To See What’s Taking Up Space On Your Windows PC
AI & Technology

How To See What’s Taking Up Space On Your Windows PC

September 4, 2026
The Tetris Company Wants Nothing To Do With The White House’s New Copycat Game
AI & Technology

The Tetris Company Wants Nothing To Do With The White House’s New Copycat Game

September 4, 2026
OpenAI Commits B to Frontline Cyber Defense, Launches MS-ISAC Pilot – Unite.AI
AI & Technology

OpenAI Commits $1B to Frontline Cyber Defense, Launches MS-ISAC Pilot – Unite.AI

September 4, 2026
Flock Cameras Are Officially Banned On State Roads In Florida
AI & Technology

Flock Cameras Are Officially Banned On State Roads In Florida

September 4, 2026
Next Post
Intel’s Bob Swan Is Out, Pat Gelsinger Is In

Intel's Bob Swan Is Out, Pat Gelsinger Is In

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Get Mortgage Pre-Approval Before You Start House Hunting

Get Mortgage Pre-Approval Before You Start House Hunting

September 4, 2026
Google DeepMind’s WeatherNext 3 Trains on Weather Station Observations to Deliver 5 km Global Forecasts, Refreshed Every Hour

Google DeepMind’s WeatherNext 3 Trains on Weather Station Observations to Deliver 5 km Global Forecasts, Refreshed Every Hour

September 4, 2026
Trump says Iran not behind Minnesota water facility hacks despite ‘hallmarks’ of Iranian meddling

Trump says Iran not behind Minnesota water facility hacks despite ‘hallmarks’ of Iranian meddling

August 31, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!