• bitcoinBitcoin(BTC)$78,600.00-0.22%
  • ethereumEthereum(ETH)$2,491.14-0.21%
  • tetherTether(USDT)$1.000.02%
  • binancecoinBNB(BNB)$739.48-2.03%
  • rippleXRP(XRP)$1.42-1.93%
  • usd-coinUSDC(USDC)$1.000.00%
  • solanaSolana(SOL)$103.11-1.01%
  • tronTRON(TRX)$0.338318-0.24%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.030.69%
  • zcashZcash(ZEC)$1,271.127.76%
  • HyperliquidHyperliquid(HYPE)$85.051.37%
  • dogecoinDogecoin(DOGE)$0.088975-1.88%
  • RainRain(RAIN)$0.016309-2.44%
  • USDSUSDS(USDS)$1.00-0.01%
  • whitebitWhiteBIT Coin(WBT)$81.251.40%
  • moneroMonero(XMR)$501.970.52%
  • chainlinkChainlink(LINK)$12.00-5.36%
  • leo-tokenLEO Token(LEO)$9.180.00%
  • cardanoCardano(ADA)$0.215909-6.46%
  • stellarStellar(XLM)$0.184322-4.76%
  • bitcoin-cashBitcoin Cash(BCH)$257.62-0.36%
  • daiDai(DAI)$1.000.01%
  • Ethena USDeEthena USDe(USDE)$1.000.00%
  • USD1USD1(USD1)$1.00-0.01%
  • litecoinLitecoin(LTC)$54.09-1.45%
  • CantonCanton(CC)$0.103366-2.45%
  • uniswapUniswap(UNI)$6.53-5.18%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.38-2.01%
  • hedera-hashgraphHedera(HBAR)$0.077887-3.66%
  • avalanche-2Avalanche(AVAX)$7.90-1.82%
  • nearNEAR Protocol(NEAR)$2.597.21%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • suiSui(SUI)$0.80-3.64%
  • shiba-inuShiba Inu(SHIB)$0.000005-3.08%
  • crypto-com-chainCronos(CRO)$0.059290-1.96%
  • paypal-usdPayPal USD(PYUSD)$1.000.00%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • tether-goldTether Gold(XAUT)$4,395.090.11%
  • MemeCoreMemeCore(M)$1.17-1.72%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • BittensorBittensor(TAO)$258.26-1.69%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • okbOKB(OKB)$113.45-1.07%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.15-0.11%
  • mantleMantle(MNT)$0.62-0.84%
  • AsterAster(ASTER)$0.74-2.46%
  • aaveAave(AAVE)$128.83-1.53%
  • polkadotPolkadot(DOT)$1.13-5.16%
  • Pump.funPump.fun(PUMP)$0.0046395.78%
  • pax-goldPAX Gold(PAXG)$4,398.520.10%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Meet FlexGen: A High-Throughput Generation Engine For Running Large Language Models (LLMs) With Limited GPU Memory

July 14, 2023
in AI & Technology
Reading Time: 5 mins read
A A
Meet FlexGen: A High-Throughput Generation Engine For Running Large Language Models (LLMs) With Limited GPU Memory
ShareShareShareShareShare

Large language models (LLMs) have recently shown impressive performance on various tasks. Generative LLM inference has never-before-seen powers, but it also faces particular difficulties. These models can include billions or trillions of parameters, meaning that running them requires tremendous memory and computing power. GPT-175B, for instance, only needs 325GB of GPU RAM to load its model weights. It would take at least five A100 (80GB) GPUs and sophisticated parallelism techniques to fit this model onto GPUs. Hence, reducing the resources needed for LLM inference has recently generated a lot of interest.

LLMs are used for various “back-of-house” operations, including benchmarking, information extraction, data wrangling, form processing, and interactive use cases like chatbots. In this study, they concentrate on a situation that they refer to as throughput-oriented generative inference. The fact that these tasks frequently call for conducting LLM inference in batches across a large number of tokens such as all the papers in a company’s corpus and are less susceptible to the delay of token generation is a significant feature of these jobs. Because of this, there are possibilities to lower resource needs in certain workloads by trading off latency for better throughput.

Three approaches have been used to reduce the resources needed for LLM inference: model compression to reduce the overall memory footprint, collaborative inference to spread out the cost of inference through decentralization, and offloading to make better use of memory on the CPU and disc. Although clear limits exist, these strategies have considerably reduced the resource needs for employing LLMs. Research in the first two methods often needs help to run 175B-scale models on a single commodity GPU because it assumes that the model fits within the GPU memory. On the other hand, due to ineffective I/O scheduling and tensor placement, cutting-edge offloading-based systems in the third category cannot reach an acceptable throughput on a single GPU.

[Sponsored] 🔥 Build your personal brand with Taplio  🚀 The 1st all-in-one AI-powered tool to grow on LinkedIn. Create better LinkedIn content 10x faster, schedule, analyze your stats & engage. Try it for free!

With a single commodity GPU, their main goal is to build effective offloading mechanisms for high-throughput generative inference. They can partially load an LLM and execute computation piecemeal by offloading it to secondary storage to operate an LLM with constrained GPU memory. The memory hierarchy is divided into three tiers in a typical system. Lower levels are slower but more plentiful, whereas higher levels are quicker but more scarce. Small batch sizes may cause bottlenecks in these systems. They may compromise latency in throughput-oriented scenarios by using a high batch size and distributing the expensive I/O operations over several memory hierarchies throughout a large batch of inputs overlapped with processing.

Even if they can compromise the delay, achieving high-throughput generative inference with constrained GPU memory is difficult. The first difficulty is coming up with a successful unloading plan. The plan should outline which tensors should be offloaded, where they should be offloaded in the three-level memory structure, and when during inference. Three types of tensors are used in generative inference: weights, activations, and key-value (KV) caching.

There are several ways to calculate because of the algorithm’s batch-by-batch, token-by-token, and layer-by-layer structure. These options come together to create a complicated design space. Offloading-based inference systems now in use inherit training-based methodologies that conduct excessive I/O and achieve throughput far below theoretical hardware constraints, making them some poor areas for inference. The creation of efficient compression algorithms presents the second problem. LLMs’ weights and activations have shown promising compression results in earlier publications. Nevertheless, when compression and offloading are coupled for high-throughput generative inference, additional compression strategies are driven by the I/O costs and memory reduction of the weights and KV cache.

Researchers from UCB, Stanford, CMU, Meta, Yandex, ETH and HSE jointly introduce FlexGen, an offloading framework for high-throughput LLM inference, to overcome these problems. FlexGen effectively schedules I/O activities, potential compression techniques, and distributed pipeline parallelism by combining memory from the GPU, CPU, and disc. These are the contributions they made:

  • They explicitly describe a search space of potential offloading options by considering the computing schedule, tensor placement, and computation delegation. They demonstrate that their search space captures a computing order with I/O complexity within 2 of optimality. Next, they create a search algorithm based on linear programming to maximize throughput within the search space.
  • They show that, without retraining or calibration, it is possible to decrease the weights and KV cache for LLMs like the OPT-175B to 4 bits with little to no accuracy loss. Fine-grained group-wise quantization, suited for lowering I/O costs and memory use during offloading, achieves this.
  • They demonstrate the efficiency of FlexGen by running OPT-175B on NVIDIA T4 (16GB) GPUs. FlexGen often permits a bigger batch size than the two cutting-edge offloading-based inference algorithms, DeepSpeed Zero-Inference and Hugging Face Accelerate. FlexGen can accomplish substantially greater throughputs as a result.

Check out the Paper and Github. All Credit For This Research Goes To the Researchers on This Project. Also, don’t forget to join our 16k+ ML SubReddit, Discord Channel, and Email Newsletter, where we share the latest AI research news, cool AI projects, and more.


YOU MAY ALSO LIKE

Why It’s Time to Abandon the ‘Set It and Forget It’ Model – Unite.AI

Lyft Is Now Offering Waymo Rides In Nashville

Aneesh Tickoo is a consulting intern at MarktechPost. He is currently pursuing his undergraduate degree in Data Science and Artificial Intelligence from the Indian Institute of Technology(IIT), Bhilai. He spends most of his time working on projects aimed at harnessing the power of machine learning. His research interest is image processing and is passionate about building solutions around it. He loves to connect with people and collaborate on interesting projects.


🔥 StoryBird.ai just dropped some amazing features. Generate an illustrated story from a prompt. Check it out here. (Sponsored)

Credit: Source link

ShareTweetSendSharePin

Related Posts

Why It’s Time to Abandon the ‘Set It and Forget It’ Model – Unite.AI
AI & Technology

Why It’s Time to Abandon the ‘Set It and Forget It’ Model – Unite.AI

September 9, 2026
Lyft Is Now Offering Waymo Rides In Nashville
AI & Technology

Lyft Is Now Offering Waymo Rides In Nashville

September 9, 2026
Harvey Secures 0M in Fresh Funding, Valuation Climbs to .5B – Unite.AI
AI & Technology

Harvey Secures $550M in Fresh Funding, Valuation Climbs to $15.5B – Unite.AI

September 9, 2026
How To Take Full Advantage Of Gemini When Planning Your Next Trip
AI & Technology

How To Take Full Advantage Of Gemini When Planning Your Next Trip

September 9, 2026
Next Post
Hims & Hers Vs. Teladoc: Different Approaches To ‘Telemedicine’ (NYSE:HIMS)

Hims & Hers Vs. Teladoc: Different Approaches To 'Telemedicine' (NYSE:HIMS)

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
The True Cost of Homeownership Goes Far Beyond Your Mortgage Payment

The True Cost of Homeownership Goes Far Beyond Your Mortgage Payment

September 2, 2026
Military firefighters battle huge Spanish wildfire

Military firefighters battle huge Spanish wildfire

September 4, 2026
Frustrations grow over police shooting in Wisconsin

Frustrations grow over police shooting in Wisconsin

September 5, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!