• bitcoinBitcoin(BTC)$78,181.00-0.31%
  • ethereumEthereum(ETH)$2,464.16-0.77%
  • tetherTether(USDT)$1.00-0.01%
  • binancecoinBNB(BNB)$721.73-4.00%
  • rippleXRP(XRP)$1.39-1.59%
  • usd-coinUSDC(USDC)$1.00-0.01%
  • solanaSolana(SOL)$101.59-1.66%
  • tronTRON(TRX)$0.338479-0.12%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.03-0.94%
  • zcashZcash(ZEC)$1,241.415.48%
  • HyperliquidHyperliquid(HYPE)$83.41-1.76%
  • dogecoinDogecoin(DOGE)$0.086086-4.29%
  • RainRain(RAIN)$0.015864-2.03%
  • USDSUSDS(USDS)$1.00-0.01%
  • moneroMonero(XMR)$510.030.98%
  • whitebitWhiteBIT Coin(WBT)$80.70-0.64%
  • chainlinkChainlink(LINK)$11.77-5.91%
  • leo-tokenLEO Token(LEO)$9.18-0.24%
  • cardanoCardano(ADA)$0.211503-3.69%
  • stellarStellar(XLM)$0.181099-3.54%
  • bitcoin-cashBitcoin Cash(BCH)$250.85-2.95%
  • daiDai(DAI)$1.000.00%
  • Ethena USDeEthena USDe(USDE)$1.00-0.01%
  • USD1USD1(USD1)$1.00-0.01%
  • litecoinLitecoin(LTC)$53.07-2.28%
  • CantonCanton(CC)$0.104015-2.97%
  • uniswapUniswap(UNI)$6.17-8.07%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.37-2.04%
  • avalanche-2Avalanche(AVAX)$7.78-2.80%
  • hedera-hashgraphHedera(HBAR)$0.076505-3.42%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • nearNEAR Protocol(NEAR)$2.487.28%
  • suiSui(SUI)$0.77-4.79%
  • shiba-inuShiba Inu(SHIB)$0.000005-3.17%
  • crypto-com-chainCronos(CRO)$0.058521-0.46%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.01%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • MemeCoreMemeCore(M)$1.20-2.04%
  • tether-goldTether Gold(XAUT)$4,391.800.77%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • BittensorBittensor(TAO)$253.68-2.46%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • okbOKB(OKB)$112.76-0.96%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.150.12%
  • mantleMantle(MNT)$0.60-5.11%
  • AsterAster(ASTER)$0.73-2.10%
  • aaveAave(AAVE)$125.31-2.56%
  • polkadotPolkadot(DOT)$1.12-9.88%
  • pax-goldPAX Gold(PAXG)$4,394.760.77%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.0564060.52%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Researchers from the University of Washington and Duke University Introduce Punica: An Artificial Intelligence System to Serve Multiple LoRA Models in a Shared GPU Cluster

November 18, 2023
in AI & Technology
Reading Time: 4 mins read
A A
Researchers from the University of Washington and Duke University Introduce Punica: An Artificial Intelligence System to Serve Multiple LoRA Models in a Shared GPU Cluster
ShareShareShareShareShare

To specialize in pre-trained large language models (LLMs) for domain-specific tasks with minimum training data, low-rank adaptation, or LoRA, is gaining popularity. Tenants may train various LoRA models at a minimal cost since LoRA greatly reduces the number of trainable parameters by keeping the pre-trained model’s weights and adding trainable rank decomposition matrices to each layer of the Transformer architecture. LoRA is now a part of several widely used fine-tuning frameworks. To meet the demands of its tenants, ML providers must thus concurrently offer many specific LoRA models. GPU resources are wasted by merely providing LoRA models as though they were individually trained. 

If k GPUs are required for every LoRA model, then k × n GPUs would appear to be needed to support n separate LoRA models. This simple method ignores the possibility of weight correlations between these LoRA models because they come from the same pre-trained models. They contend that an effective system that supports several distinct LoRA models must adhere to three design principles. Since (G1) GPUs are costly and in short supply, multi-tenant LoRA serving workloads must be concentrated onto a small number of GPUs to maximize GPU usage. (G2) Batching is one of the best, if not the best, ways to combine ML workloads to increase performance and GPU usage, as previous studies have noted. But they are batching only functions in cases where requests are made for identical models. As a result, they must allow batching for various LoRA models. (G3) Most model serving costs are attributed to the decode stage. So, all they have to concentrate on is the amazing stage performance. They can use simple methods, such as on-demand loading of LoRA model weights, for other less crucial components of the model serving. Based on these three criteria, researchers from the University of Washington and Duke University developed and built Punica, a multi-tenant serving framework for LoRA models on a shared GPU cluster. Segmented Gather Matrix-Vector Multiplication (SGMV), a new CUDA kernel, is one of the main innovations. 

Batching GPU operations for the simultaneous execution of several distinct LoRA models is made possible by SGMV. By reducing the number of copies of the pre-trained model that a GPU must keep in memory, SGMV dramatically increases GPU efficiency in both memory and computation. They combine several cutting-edge methods for system optimization with this new CUDA kernel. Surprisingly, they find very few performance differences when batching the same LoRA models versus batching different LoRA models. SGMV permits batching requests from several LoRA models. Simultaneously, the delay of the LoRA model on-demand loading is mere milliseconds. 

Punica may now condense user requests to a smaller group of GPUs without being limited by the LoRA models currently executing on the GPUs. Punica uses the following two methods to arrange tasks for several tenants. Punica directs a fresh request to a select group of GPUs currently in use, ensuring they are utilized to their maximum potential. Punica will only commit further GPU resources once the current GPUs are completely used. Punica moves active requests for consolidation regularly. This makes it possible to release GPU resources that Punica has been assigned. On NVIDIA A100 GPU clusters, they assess LoRA models derived from the Llama2 7B, 13B, and 70B models.

Punica adds a 2ms delay per token and delivers 12x greater throughput than state-of-the-art LLM serving solutions with the same GPU resources. The following are the contributions made by this paper: 

• They recognize the potential for batch-processing requests from various LoRA models. 

• They create and put into practice a CUDA kernel that is effective for running many LoRA models at once. • They provide innovative scheduling techniques to combine tasks from many tenants in LoRA.


Check out the Paper and Github. All credit for this research goes to the researchers of this project. Also, don’t forget to join our 33k+ ML SubReddit, 41k+ Facebook Community, Discord Channel, and Email Newsletter, where we share the latest AI research news, cool AI projects, and more.

If you like our work, you will love our newsletter..


YOU MAY ALSO LIKE

Apple Wallet Is Not The Same As Apple Pay: Here’s How They Differ

Google Open-Sources Mantis: A Modular Skills Toolkit That Lets Coding Agents Find, Reproduce and Patch Vulnerabilities

Aneesh Tickoo is a consulting intern at MarktechPost. He is currently pursuing his undergraduate degree in Data Science and Artificial Intelligence from the Indian Institute of Technology(IIT), Bhilai. He spends most of his time working on projects aimed at harnessing the power of machine learning. His research interest is image processing and is passionate about building solutions around it. He loves to connect with people and collaborate on interesting projects.


🔥 Join The AI Startup Newsletter To Learn About Latest AI Startups

Credit: Source link

ShareTweetSendSharePin

Related Posts

Apple Wallet Is Not The Same As Apple Pay: Here’s How They Differ
AI & Technology

Apple Wallet Is Not The Same As Apple Pay: Here’s How They Differ

September 9, 2026
Google Open-Sources Mantis: A Modular Skills Toolkit That Lets Coding Agents Find, Reproduce and Patch Vulnerabilities
AI & Technology

Google Open-Sources Mantis: A Modular Skills Toolkit That Lets Coding Agents Find, Reproduce and Patch Vulnerabilities

September 9, 2026
Blizzard Employees Have Ratified Their First Union Contracts
AI & Technology

Blizzard Employees Have Ratified Their First Union Contracts

September 9, 2026
OpenAI Names Paul Christiano to Foundation Board and Safety Committee – Unite.AI
AI & Technology

OpenAI Names Paul Christiano to Foundation Board and Safety Committee – Unite.AI

September 9, 2026
Next Post
Meet the Press NOW – Nov. 3

Meet the Press NOW – Nov. 3

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Watch Senator Lindsey Graham funeral service at the U.S. Capitol | NBC News

Watch Senator Lindsey Graham funeral service at the U.S. Capitol | NBC News

September 4, 2026
What Is Model Routing? How AI Systems Choose the Right Model for Every Request – Unite.AI

What Is Model Routing? How AI Systems Choose the Right Model for Every Request – Unite.AI

September 6, 2026
Canada’s retaliatory US tariffs set to take effect as trade dispute grows – The Guardian

Canada’s retaliatory US tariffs set to take effect as trade dispute grows – The Guardian

September 7, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!