• bitcoinBitcoin(BTC)$78,432.00-0.33%
  • ethereumEthereum(ETH)$2,476.29-0.64%
  • tetherTether(USDT)$1.00-0.01%
  • binancecoinBNB(BNB)$725.52-3.37%
  • rippleXRP(XRP)$1.39-1.56%
  • usd-coinUSDC(USDC)$1.00-0.01%
  • solanaSolana(SOL)$101.95-1.40%
  • tronTRON(TRX)$0.3395900.19%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.03-0.94%
  • zcashZcash(ZEC)$1,244.615.13%
  • HyperliquidHyperliquid(HYPE)$84.19-1.51%
  • dogecoinDogecoin(DOGE)$0.086081-4.15%
  • RainRain(RAIN)$0.0163121.88%
  • USDSUSDS(USDS)$1.00-0.01%
  • moneroMonero(XMR)$513.992.78%
  • whitebitWhiteBIT Coin(WBT)$80.95-0.71%
  • chainlinkChainlink(LINK)$11.85-4.60%
  • leo-tokenLEO Token(LEO)$9.20-0.31%
  • cardanoCardano(ADA)$0.213363-1.93%
  • stellarStellar(XLM)$0.180587-3.72%
  • bitcoin-cashBitcoin Cash(BCH)$252.30-2.16%
  • daiDai(DAI)$1.000.01%
  • Ethena USDeEthena USDe(USDE)$1.00-0.02%
  • USD1USD1(USD1)$1.00-0.01%
  • CantonCanton(CC)$0.104145-4.37%
  • litecoinLitecoin(LTC)$52.96-2.03%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.38-0.87%
  • uniswapUniswap(UNI)$6.05-10.59%
  • hedera-hashgraphHedera(HBAR)$0.077072-2.02%
  • avalanche-2Avalanche(AVAX)$7.81-1.95%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • nearNEAR Protocol(NEAR)$2.5210.23%
  • suiSui(SUI)$0.77-5.00%
  • shiba-inuShiba Inu(SHIB)$0.000005-2.50%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.02%
  • crypto-com-chainCronos(CRO)$0.058141-3.61%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • MemeCoreMemeCore(M)$1.231.78%
  • tether-goldTether Gold(XAUT)$4,412.020.75%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • BittensorBittensor(TAO)$254.29-0.50%
  • Ripple USDRipple USD(RLUSD)$1.00-0.01%
  • okbOKB(OKB)$113.20-1.19%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.150.00%
  • mantleMantle(MNT)$0.60-4.71%
  • AsterAster(ASTER)$0.73-3.72%
  • aaveAave(AAVE)$125.43-2.42%
  • pax-goldPAX Gold(PAXG)$4,415.530.75%
  • polkadotPolkadot(DOT)$1.12-7.44%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.0566781.25%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

A Team of UC Berkeley and Stanford Researchers Introduce S-LoRA: An Artificial Intelligence System Designed for the Scalable Serving of Many LoRA Adapters

November 12, 2023
in AI & Technology
Reading Time: 4 mins read
A A
A Team of UC Berkeley and Stanford Researchers Introduce S-LoRA: An Artificial Intelligence System Designed for the Scalable Serving of Many LoRA Adapters
ShareShareShareShareShare

A team of UC Berkeley and Stanford researchers have developed a new parameter-efficient fine-tuning method called Low-Rank Adaptation (LoRA) for deploying LLMs. S-LoRA was designed to enable the efficient deployment of many LoRA adapters. S-LoRA allows thousands of adapters to run on a single GPU or across multiple GPUs with minimal overhead. The method introduces unified paging to optimize GPU memory usage, utilizing novel tensor parallelism and custom CUDA kernels for heterogeneous batch processing. These techniques significantly reduce the computational requirements for deploying LLMs in real-world applications.

LoRA is a highly efficient fine-tuning technique for customizing pre-trained LLMs to new tasks, dramatically reducing the trainable parameters while maintaining high accuracy. LoRA is widely embraced, resulting in the creation of countless LoRA adapters for LLMs and diffusion models. In today’s applications, LLMs are pervasive, catering to various domains and tasks.

Modern applications extensively utilize LLMs, and the pretrain-then-finetune method has resulted in the creation of multiple fine-tuned versions of a single base LLM, each customized for specific tasks or domains. LoRA is a parameter-efficient fine-tuning technique that tailors pre-trained LLMs for new tasks, significantly decreasing the number of trainable parameters while maintaining high accuracy.

S-LoRA leverages LoRA to efficiently fine-tune a base model for a wide range of tasks, generating a substantial collection of LoRA adapters from a single model. It introduces Unified Paging, which optimizes GPU memory usage by managing dynamic adapter weights and KV cache tensors within a unified memory pool. S-LoRA enables the serving of thousands of LoRA adapters with minimal overhead. The approach can enhance throughput fourfold and significantly scale up the number of supported adapters compared to leading libraries like HuggingFace PEFT and vLLM.

S-LoRA efficiently handles 2,000 adapters simultaneously with minimal overhead, maintaining low computational costs. It outperforms vLLM-packed by up to 4 times for a few adapters and up to 30 times over PEFT while accommodating a significantly larger adapter count. S-LoRA surpasses its variations, S-LoRA-bmm and S-LoRA-no-unifymem, in throughput and latency, highlighting the effectiveness of memory pooling and custom kernels. The system’s scalability is primarily limited by available main memory, demonstrating robust performance for real-world workloads. S-LoRA’s impressive capabilities make it a powerful solution for adapting large language models to various tasks.

The research aims to enhance performance by investigating optimization avenues such as quantization, sparsification, and refining model architectures. It explores the implementation of decomposed computation techniques for both the base model and adapters, along with the development of custom CUDA kernels for enhanced support. The focus also extends to addressing auto-regressive features and parameter-efficient adapters within LLM serving, seeking to identify and bridge optimization gaps in current model serving systems.

In conclusion, S-LoRA has introduced unified paging to combat memory fragmentation, leading to increased batch sizes and improved scalability in serving. The study presents a scalable LoRA serving solution, addressing the previously unexplored challenge of serving fine-tuned variants at scale. The work optimizes LoRA serving through algorithmic techniques like quantization, sparsification, and model architecture enhancements, complementing system-level improvements.


Check out the Paper and Github. All credit for this research goes to the researchers of this project. Also, don’t forget to join our 32k+ ML SubReddit, 41k+ Facebook Community, Discord Channel, and Email Newsletter, where we share the latest AI research news, cool AI projects, and more.

If you like our work, you will love our newsletter..

We are also on Telegram and WhatsApp.


YOU MAY ALSO LIKE

Apple Wallet Is Not The Same As Apple Pay: Here’s How They Differ

Google Open-Sources Mantis: A Modular Skills Toolkit That Lets Coding Agents Find, Reproduce and Patch Vulnerabilities

Hello, My name is Adnan Hassan. I am a consulting intern at Marktechpost and soon to be a management trainee at American Express. I am currently pursuing a dual degree at the Indian Institute of Technology, Kharagpur. I am passionate about technology and want to create new products that make a difference.


🔥 Meet Retouch4me: A Family of Artificial Intelligence-Powered Plug-Ins for Photography Retouching

Credit: Source link

ShareTweetSendSharePin

Related Posts

Apple Wallet Is Not The Same As Apple Pay: Here’s How They Differ
AI & Technology

Apple Wallet Is Not The Same As Apple Pay: Here’s How They Differ

September 9, 2026
Google Open-Sources Mantis: A Modular Skills Toolkit That Lets Coding Agents Find, Reproduce and Patch Vulnerabilities
AI & Technology

Google Open-Sources Mantis: A Modular Skills Toolkit That Lets Coding Agents Find, Reproduce and Patch Vulnerabilities

September 9, 2026
Muse, The Band, Lost Its Social Media Handles To Muse, Meta’s New AI Agent
AI & Technology

Muse, The Band, Lost Its Social Media Handles To Muse, Meta’s New AI Agent

September 9, 2026
Blizzard Employees Have Ratified Their First Union Contracts
AI & Technology

Blizzard Employees Have Ratified Their First Union Contracts

September 9, 2026
Next Post
This Morning’s Top Headlines – June 23 | Morning News NOW

This Morning’s Top Headlines – June 23 | Morning News NOW

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Nvidia Makes MediaTek Partnership Even Bigger, Huang Says

Nvidia Makes MediaTek Partnership Even Bigger, Huang Says

September 4, 2026
5 dead, 5 injured after Amazon cargo plane overruns runway at MIA, sheriff says – NBC 6 South Florida

5 dead, 5 injured after Amazon cargo plane overruns runway at MIA, sheriff says – NBC 6 South Florida

September 7, 2026
NBC Nightly News with Tom Llamas Full Episode – July 28

NBC Nightly News with Tom Llamas Full Episode – July 28

September 3, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!