• bitcoinBitcoin(BTC)$78,025.000.23%
  • ethereumEthereum(ETH)$2,452.230.45%
  • tetherTether(USDT)$1.000.00%
  • binancecoinBNB(BNB)$691.47-0.37%
  • rippleXRP(XRP)$1.390.19%
  • usd-coinUSDC(USDC)$1.000.00%
  • solanaSolana(SOL)$105.090.18%
  • tronTRON(TRX)$0.338063-0.28%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.040.52%
  • HyperliquidHyperliquid(HYPE)$83.252.13%
  • zcashZcash(ZEC)$837.144.03%
  • dogecoinDogecoin(DOGE)$0.0852950.01%
  • RainRain(RAIN)$0.017638-0.07%
  • USDSUSDS(USDS)$1.000.02%
  • leo-tokenLEO Token(LEO)$9.632.20%
  • moneroMonero(XMR)$464.55-0.71%
  • chainlinkChainlink(LINK)$11.43-0.27%
  • whitebitWhiteBIT Coin(WBT)$71.960.29%
  • cardanoCardano(ADA)$0.200837-1.02%
  • stellarStellar(XLM)$0.1790800.09%
  • bitcoin-cashBitcoin Cash(BCH)$245.84-0.61%
  • CantonCanton(CC)$0.1169916.47%
  • daiDai(DAI)$1.000.00%
  • USD1USD1(USD1)$1.00-0.01%
  • Ethena USDeEthena USDe(USDE)$1.000.02%
  • litecoinLitecoin(LTC)$48.930.06%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.36-0.96%
  • Global DollarGlobal Dollar(USDG)$1.000.03%
  • hedera-hashgraphHedera(HBAR)$0.075489-0.44%
  • avalanche-2Avalanche(AVAX)$7.310.22%
  • suiSui(SUI)$0.75-0.72%
  • shiba-inuShiba Inu(SHIB)$0.000005-0.97%
  • uniswapUniswap(UNI)$4.624.00%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • crypto-com-chainCronos(CRO)$0.0578670.89%
  • paypal-usdPayPal USD(PYUSD)$1.000.00%
  • Circle USYCCircle USYC(USYC)$1.140.00%
  • tether-goldTether Gold(XAUT)$4,456.53-0.27%
  • nearNEAR Protocol(NEAR)$1.841.02%
  • okbOKB(OKB)$112.662.15%
  • MemeCoreMemeCore(M)$1.046.97%
  • Ripple USDRipple USD(RLUSD)$1.000.04%
  • BittensorBittensor(TAO)$238.23-0.38%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.14-0.09%
  • pax-goldPAX Gold(PAXG)$4,461.74-0.26%
  • Pump.funPump.fun(PUMP)$0.0048293.29%
  • aaveAave(AAVE)$123.320.37%
  • AsterAster(ASTER)$0.702.17%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.0578800.33%
  • OndoOndo(ONDO)$0.353074-0.90%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Moonshot AI Open-Sources MoonEP: A Perfectly Balanced Expert Parallelism Library for MoE Training

July 30, 2026
in AI & Technology
Reading Time: 14 mins read
A A
Moonshot AI Open-Sources MoonEP: A Perfectly Balanced Expert Parallelism Library for MoE Training
ShareShareShareShareShare

Moonshot AI has open-sourced MoonEP, an Expert Parallelism (EP) communication library for distributed Mixture-of-Experts (MoE) workloads. The team announced the release as a library built to make expert-parallel communication more efficient at scale. It ships under an MIT license.

MoonEP arrived as part of Kimi K3 Open Day. Alongside the K3 model weights and technical report, Moonshot released three infrastructure codebases: MoonEP, FlashKDA, and AgentEnv. FlashKDA had already been open-sourced; MoonEP and AgentEnv were published with this release. MoonEP is one of the innovations behind a claimed 2.5× improvement in scaling efficiency for Kimi K3, a 2.8-trillion-parameter MoE model with native vision and a 1M-token context window.

The problem MoonEP targets

In expert parallelism, a router sends each token to its top-K experts, which live on different ranks. Routers are rarely balanced. Some experts get far more tokens than others.

The repository quantifies skew with maxvio, defined as max_e (T_e / T̄) − 1, where T_e is tokens routed to expert e and T̄ is the expected count under perfect balance. A maxvio of 0 means perfectly balanced.

Imbalance costs are structural, not incidental. A collective’s latency is set by its slowest participant, so the hottest rank determines iteration time. Worse, token counts per rank change every step. Those dynamic activation shapes fragment GPU memory and force per-layer host synchronization.

The core idea: dynamic redundant experts

MoonEP’s main mention is a hard invariant. Every rank receives exactly S × K tokens, no matter how skewed the routing is — where S is input tokens per rank and K is routed top-k per token.

It achieves this by planning a small number of redundant experts online, directly from the current router outputs. Those duplicated experts are prefetched before expert computation. In the backward pass, their gradients are reduced back to their home ranks.

The design is set up into three properties:

  1. Perfect balance: the S × K guarantee above, via online-planned redundant experts.
  2. Online planning: a near-optimal GPU planning kernel with negligible overhead. It is implemented in the CUTLASS CuTe DSL; setup.py pins nvidia-cutlass-dsl==4.4.2.
  3. Zero copy and static shapes: fused permute/unpermute. Tokens are written directly into their expert-grouped positions on remote ranks, and buffer views are returned to the computation. Only a fixed S × K buffer is needed, and statically known shapes eliminate per-layer MoE host synchronization.

The interactive explainer below computes the resulting buffer and prefetch-pool sizes live from a config you control.

The memory contract

MoonEP’s contract with a training or inference framework is specific: one contiguous symmetric-memory weight tensor per expert projection, plus a planner-produced cu_seqlens. The VM group GEMM consumes a single [E+B, H, H'] weight tensor, where E is total routed experts, B is prefetch slots per rank, H is hidden size, and H' is the expert FFN intermediate size. The cu_seqlens[E+B] returned by dispatch selects which expert rows are active.

Contiguity is a hard requirement, because the group GEMM addresses experts purely by row index. The layout splits cleanly:

YOU MAY ALSO LIKE

Sony and Warner Chappell Sue Anthropic Over Claude Lyric Training – Unite.AI

Don’t Fall For This Digital TV Antenna Myth

  • Rows [0, E) hold all ranks’ local experts, E/R rows per rank. Each chunk physically is the home rank’s parameter memory, mapped everywhere via symmetric memory.
  • Rows [E, E+B) are local prefetch slots, filled by buffer.prefetch_weight.

The prefetch slots draw from a process-global pool shared by all layers. That detail matters: the extra memory cost is B expert weights per projection in total, not per layer.

How you set B depends on the workload. Training must use B = E/R, because the planner duplicates experts from at most one remote home group per rank. That bound guarantees every expert the group GEMM touches is local. Inference allows B < E/R, and the README recommends B = 3–4. If a rank needs more distinct remote experts than B, the group GEMM reads overflow weights straight from the home rank through the symmetric mapping — slightly slower, with no impact on correctness.

Training mirrors the weight layout in fp32 with a [E+B, H, H'] grad buffer per projection. Critically, rows [E, E+B) are backed by a separate reduce buffer, not by the parameter grads. Duplicated experts’ gradients are temporary and must stay invisible to the framework’s own grad reduce. Each rank maps all R reduce buffers as one [R, B, H, H'] view, then reduce_grad reads its own experts’ slots from every rank over NVLink, accumulates into the local parameter grad, and zeroes consumed slots.

Benchmarks against DeepEP v2

Both published benchmarks run on H20 with EP=8, sweeping router imbalance. The comparison script benchmarks/bench_vs_deepep.py defaults to S=8192, E=384, H=7168, K=8, H'=2048, and 32 SMs, with maxvio targets of 0.2, 1, 10, and 20. Both libraries receive an identical routing matrix from a shared seed.

Three findings were reported on their github page. Zero copy makes raw communication faster by eliminating the comm-buffer → user-buffer copy that dominates the epilogue, so MoonEP’s comm time sits consistently below DeepEP v2 at every imbalance level. Perfect balance makes MoonEP nearly immune to skew: its comm time stays almost flat as maxvio grows, while DeepEP v2 — whose latency is set by the hottest rank — degrades steadily.

Key Takeaways

  • MoonEP claims every EP rank receives exactly S × K tokens, regardless of router skew.
  • Balance comes from redundant experts planned online on-GPU, then prefetched before expert compute.
  • Static shapes remove per-layer MoE host sync and stop the memory fragmentation that OOMs DeepEP.
  • Training requires B = E/R prefetch slots; inference can drop to B = 3–4 with no correctness cost.
  • Released MIT under Kimi K3 Open Day, alongside FlashKDA and AgentEnv.

Sources: MoonshotAI/MoonEP on GitHub and @Kimi_Moonshot announcement


Michal Sutter is a data science professional with a Master of Science in Data Science from the University of Padova. With a solid foundation in statistical analysis, machine learning, and data engineering, Michal excels at transforming complex datasets into actionable insights.

Credit: Source link

ShareTweetSendSharePin

Related Posts

Sony and Warner Chappell Sue Anthropic Over Claude Lyric Training – Unite.AI
AI & Technology

Sony and Warner Chappell Sue Anthropic Over Claude Lyric Training – Unite.AI

August 29, 2026
Don’t Fall For This Digital TV Antenna Myth
AI & Technology

Don’t Fall For This Digital TV Antenna Myth

August 29, 2026
Google AI Releases Gemini Omni 1.1 Flash: 40-Second Scene Extension, First/Last Frame Control, and 4K Upscaling
AI & Technology

Google AI Releases Gemini Omni 1.1 Flash: 40-Second Scene Extension, First/Last Frame Control, and 4K Upscaling

August 29, 2026
What You Should Look For In A USB To Bluetooth Adapter For PC
AI & Technology

What You Should Look For In A USB To Bluetooth Adapter For PC

August 29, 2026
Next Post
U.S. launches second night of strikes on Iran

U.S. launches second night of strikes on Iran

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
U.S. economy lost 23,000 jobs in July

U.S. economy lost 23,000 jobs in July

August 28, 2026
Police arrest armed suspect at Trump golf course

Police arrest armed suspect at Trump golf course

August 29, 2026
Suspected burglar uses plant to hide face from ring camera

Suspected burglar uses plant to hide face from ring camera

August 28, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!