• bitcoinBitcoin(BTC)$84,417.00-2.04%
  • ethereumEthereum(ETH)$2,683.01-2.55%
  • tetherTether(USDT)$1.000.00%
  • binancecoinBNB(BNB)$767.12-2.68%
  • rippleXRP(XRP)$1.50-4.66%
  • usd-coinUSDC(USDC)$1.00-0.01%
  • solanaSolana(SOL)$114.88-3.01%
  • tronTRON(TRX)$0.341860-0.04%
  • zcashZcash(ZEC)$1,498.65-7.74%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.03-0.49%
  • HyperliquidHyperliquid(HYPE)$93.89-3.30%
  • dogecoinDogecoin(DOGE)$0.092542-7.71%
  • moneroMonero(XMR)$552.94-3.51%
  • whitebitWhiteBIT Coin(WBT)$84.70-2.28%
  • USDSUSDS(USDS)$1.000.00%
  • chainlinkChainlink(LINK)$12.35-5.18%
  • cardanoCardano(ADA)$0.238407-6.13%
  • RainRain(RAIN)$0.012243-6.54%
  • leo-tokenLEO Token(LEO)$8.990.37%
  • stellarStellar(XLM)$0.202036-6.46%
  • bitcoin-cashBitcoin Cash(BCH)$337.89-2.07%
  • uniswapUniswap(UNI)$9.27-9.44%
  • nearNEAR Protocol(NEAR)$4.29-2.39%
  • Ethena USDeEthena USDe(USDE)$1.00-0.01%
  • litecoinLitecoin(LTC)$61.66-1.90%
  • daiDai(DAI)$1.000.01%
  • avalanche-2Avalanche(AVAX)$10.27-8.62%
  • USD1USD1(USD1)$1.000.00%
  • CantonCanton(CC)$0.109960-4.28%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.42-3.93%
  • hedera-hashgraphHedera(HBAR)$0.090431-8.72%
  • suiSui(SUI)$0.96-6.37%
  • shiba-inuShiba Inu(SHIB)$0.000006-7.54%
  • BittensorBittensor(TAO)$287.12-8.76%
  • Global DollarGlobal Dollar(USDG)$1.00-0.02%
  • crypto-com-chainCronos(CRO)$0.061460-7.76%
  • BitwayBitway(BTW)$1.0315.76%
  • MemeCoreMemeCore(M)$1.23-6.63%
  • paypal-usdPayPal USD(PYUSD)$1.000.00%
  • tether-goldTether Gold(XAUT)$4,289.07-1.54%
  • okbOKB(OKB)$118.97-3.13%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.15-0.03%
  • mantleMantle(MNT)$0.65-2.46%
  • aaveAave(AAVE)$138.95-5.94%
  • EthenaEthena(ENA)$0.205440-5.35%
  • OndoOndo(ONDO)$0.412363-6.42%
  • AsterAster(ASTER)$0.69-4.79%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

ZenFlow: A New DeepSpeed Extension Designed as a Stall-Free Offloading Engine for Large Language Model (LLM) Training

August 20, 2025
in AI & Technology
Reading Time: 6 mins read
A A
ZenFlow: A New DeepSpeed Extension Designed as a Stall-Free Offloading Engine for Large Language Model (LLM) Training
ShareShareShareShareShare

The DeepSpeed team unveiled ZenFlow, a new offloading engine designed to overcome a major bottleneck in large language model (LLM) training: CPU-induced GPU stalls. While offloading optimizers and gradients to CPU memory reduces GPU memory pressure, traditional frameworks like ZeRO-Offload and ZeRO-Infinity often leave expensive GPUs idle for most of each training step—waiting on slow CPU updates and PCIe transfers. For example, fine-tuning Llama 2-7B on 4× A100 GPUs with full offloading can balloon step time from 0.5s to over 7s, a 14× slowdown. ZenFlow eliminates these stalls by decoupling GPU and CPU computation with importance-aware pipelining, delivering up to 5× end-to-end speedup over ZeRO-Offload and reducing GPU stalls by more than 85%.

How ZenFlow Works

  • Importance-Aware Gradient Updates: ZenFlow prioritizes the top-k most impactful gradients for immediate GPU updates, while deferring less important gradients to asynchronous CPU-side accumulation. This reduces per-step gradient traffic by nearly 50% and PCIe bandwidth pressure by about 2× compared to ZeRO-Offload.
  • Bounded-Asynchronous CPU Accumulation: Non-critical gradients are batched and updated asynchronously on the CPU, hiding CPU work behind GPU compute. This ensures GPUs are always busy, avoiding stalls and maximizing hardware utilization.
  • Lightweight Gradient Selection: ZenFlow replaces full gradient AllGather with a lightweight, per-column gradient norm proxy, reducing communication volume by over 4,000× with minimal impact on accuracy. This enables efficient scaling across multi-GPU clusters.
  • Zero Code Changes, Minimal Configuration: ZenFlow is built into DeepSpeed and requires only minor JSON configuration changes. Users set parameters like topk_ratio (e.g., 0.05 for top 5% of gradients) and enable adaptive strategies with select_strategy, select_interval, and update_interval set to "auto".
  • Auto-Tuned Performance: The engine adapts update intervals at runtime, eliminating the need for manual tuning and ensuring maximum efficiency as training dynamics evolve.
https://arxiv.org/abs/2505.12242

Performance Highlights

Feature Impact
Up to 5× end-to-end speedup Faster convergence, lower costs
>85% reduction in GPU stalls Higher GPU utilization
≈2× lower PCIe traffic Less cluster bandwidth pressure
No accuracy loss on GLUE benchmarks Maintains model quality
Lightweight gradient selection Scales efficiently to multi-GPU clusters
Auto-tuning No manual parameter tuning required

Practical Usage

Integration: ZenFlow is a drop-in extension for DeepSpeed’s ZeRO-Offload. No code changes are needed; only configuration updates in the DeepSpeed JSON file are required.

YOU MAY ALSO LIKE

Meta Brings FDA-Cleared Hearing Enhancement To Its Smart Glasses

Microsoft’s New Surface Pro 12 And Surface Laptop 13 Feature Snapdragon X2 Plus Chips

Example Use Case: The DeepSpeedExamples repository includes a ZenFlow finetuning example on the GLUE benchmark. Users can run this with a simple script (bash finetune_gpt_glue.sh), following setup and configuration instructions in the repo’s README. The example demonstrates CPU optimizer offload with ZenFlow asynchronous updates, providing a practical starting point for experimentation.

Configuration Example:

Copy CodeCopiedUse a different Browser
"zero_optimization": {
  "stage": 2,
  "offload_optimizer": {
    "device": "cpu",
    "pin_memory": true
  },
  "zenflow": {
    "topk_ratio": 0.05,
    "select_strategy": "auto",
    "select_interval": "auto",
    "update_interval": 4,
    "full_warm_up_rounds": 0,
    "overlap_step": true
  }
}

Getting Started: Refer to the DeepSpeed-ZenFlow finetuning example and the official tutorial for step-by-step guidance.

Summary

ZenFlow is a significant leap forward for anyone training or fine-tuning large language models on limited GPU resources. By effectively eliminating CPU-induced GPU stalls, it unlocks higher throughput and lower total cost of training, without sacrificing model accuracy. The approach is particularly valuable for organizations scaling LLM workloads across heterogeneous hardware or seeking to maximize GPU utilization in cloud or on-prem clusters.

For technical teams, the combination of automatic tuning, minimal configuration, and seamless integration with DeepSpeed makes ZenFlow both accessible and powerful. The provided examples and documentation lower the barrier to adoption, enabling rapid experimentation and deployment.

ZenFlow redefines offloading for LLM training, delivering stall-free, high-throughput fine-tuning with minimal configuration overhead—a must-try for anyone pushing the boundaries of large-scale AI.


Check out the Technical Paper, GitHub Page and Blog. Feel free to check out our GitHub Page for Tutorials, Codes and Notebooks. Also, feel free to follow us on Twitter and don’t forget to join our 100k+ ML SubReddit and Subscribe to our Newsletter.

The post ZenFlow: A New DeepSpeed Extension Designed as a Stall-Free Offloading Engine for Large Language Model (LLM) Training appeared first on MarkTechPost.

Credit: Source link

ShareTweetSendSharePin

Related Posts

Meta Brings FDA-Cleared Hearing Enhancement To Its Smart Glasses
AI & Technology

Meta Brings FDA-Cleared Hearing Enhancement To Its Smart Glasses

September 23, 2026
Microsoft’s New Surface Pro 12 And Surface Laptop 13 Feature Snapdragon X2 Plus Chips
AI & Technology

Microsoft’s New Surface Pro 12 And Surface Laptop 13 Feature Snapdragon X2 Plus Chips

September 23, 2026
Google Releases Gemini 3.8 Flash TTS and Flash-Lite TTS With Prompt-Based Voice Design
AI & Technology

Google Releases Gemini 3.8 Flash TTS and Flash-Lite TTS With Prompt-Based Voice Design

September 23, 2026
NVIDIA Releases Nemotron 3 Diarization: A 100M-Parameter Open-Weight Model That Tracks 8 Speakers in Real Time
AI & Technology

NVIDIA Releases Nemotron 3 Diarization: A 100M-Parameter Open-Weight Model That Tracks 8 Speakers in Real Time

September 23, 2026
Next Post
Synthesia: The AI Avatar Generator Rethinking Corporate Communication

Synthesia: The AI Avatar Generator Rethinking Corporate Communication

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
I Bought A Business On Credit Cards (It’s Not Making Money)

I Bought A Business On Credit Cards (It’s Not Making Money)

September 22, 2026
Cheesecake Factory: Traffic Is Back At The Table

Cheesecake Factory: Traffic Is Back At The Table

September 21, 2026
NEXT plc 2027 Q2 – Results – Earnings Call Presentation (OTCMKTS:NXGPY) 2026-09-18

NEXT plc 2027 Q2 – Results – Earnings Call Presentation (OTCMKTS:NXGPY) 2026-09-18

September 18, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!