• bitcoinBitcoin(BTC)$64,271.00-0.40%
  • ethereumEthereum(ETH)$1,901.63-0.10%
  • tetherTether(USDT)$1.000.00%
  • binancecoinBNB(BNB)$592.370.00%
  • usd-coinUSDC(USDC)$1.000.00%
  • rippleXRP(XRP)$1.04-2.50%
  • solanaSolana(SOL)$72.60-1.70%
  • tronTRON(TRX)$0.326795-0.10%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.02-1.40%
  • HyperliquidHyperliquid(HYPE)$55.78-1.00%
  • dogecoinDogecoin(DOGE)$0.069207-1.30%
  • USDSUSDS(USDS)$1.000.00%
  • RainRain(RAIN)$0.012539-0.30%
  • leo-tokenLEO Token(LEO)$9.750.00%
  • zcashZcash(ZEC)$503.66-2.40%
  • cardanoCardano(ADA)$0.2019005.00%
  • moneroMonero(XMR)$369.701.20%
  • whitebitWhiteBIT Coin(WBT)$55.65-0.50%
  • chainlinkChainlink(LINK)$8.190.60%
  • stellarStellar(XLM)$0.161620-2.30%
  • daiDai(DAI)$1.000.00%
  • bitcoin-cashBitcoin Cash(BCH)$212.74-0.70%
  • USD1USD1(USD1)$1.000.00%
  • Ethena USDeEthena USDe(USDE)$1.000.00%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.37-2.20%
  • CantonCanton(CC)$0.090813-11.90%
  • litecoinLitecoin(LTC)$45.510.50%
  • Global DollarGlobal Dollar(USDG)$1.00-0.10%
  • Circle USYCCircle USYC(USYC)$1.130.00%
  • hedera-hashgraphHedera(HBAR)$0.068293-1.40%
  • avalanche-2Avalanche(AVAX)$6.42-3.10%
  • paypal-usdPayPal USD(PYUSD)$1.000.00%
  • shiba-inuShiba Inu(SHIB)$0.000005-4.70%
  • suiSui(SUI)$0.67-2.50%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • tether-goldTether Gold(XAUT)$4,217.540.10%
  • uniswapUniswap(UNI)$4.06-0.10%
  • crypto-com-chainCronos(CRO)$0.053402-1.20%
  • nearNEAR Protocol(NEAR)$1.66-2.10%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.140.20%
  • pax-goldPAX Gold(PAXG)$4,228.330.10%
  • BittensorBittensor(TAO)$191.96-2.70%
  • okbOKB(OKB)$85.59-0.50%
  • OndoOndo(ONDO)$0.355229-3.70%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.052595-1.30%
  • HTX DAOHTX DAO(HTX)$0.000002-0.10%
  • AsterAster(ASTER)$0.60-0.80%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • usddUSDD(USDD)$1.000.10%
  • MemeCoreMemeCore(M)$1.13-5.30%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

NVIDIA AI Releases Molt: A PyTorch-Native Agentic Reinforcement Learning Framework

August 2, 2026
in AI & Technology
Reading Time: 20 mins read
A A
NVIDIA AI Releases Molt: A PyTorch-Native Agentic Reinforcement Learning Framework
ShareShareShareShareShare

Agentic reinforcement learning research is constant algorithm modification. New estimators, new pipeline stages, new rollout schemes. In mainstream frameworks each change threads through layers of trainer, distributed backend, and rollout glue. That cost lands on the researcher at every iteration.

Molt, from NVIDIA’s NeMo team, targets that cost directly. Its a PyTorch-native agentic RL framework with an unusual design target. The codebase should be compact enough for a researcher to hold in their head, and for an AI coding assistant to read and reason about in its entirety. The stated footprint is roughly 8.6K lines of RL code, measured by tracing the import graph from each framework’s RL entry point. The same method counts about 62K lines for verl, 25K for slime, and 7.2K for OpenRLHF.

Is it deployable?

Yes. Molt ships under Apache 2.0 with launch codes, Slurm scripts, and a prebuilt container. But the research paper positions it as research infrastructure, not a production training service, and hardware is the real gate. The shipped recipes assume 2 nodes of 8 H100 GPUs, split 8 for training and 8 for rollout.

That puts Molt in reach of frontier and frontier-adjacent labs, well-funded AI startups doing post-training, enterprise AI research groups in finance, healthcare, and robotics that train agents against proprietary environments, and academic labs with multi-node H100/H200 access. Applications include multi-turn tool-use agents, code-execution agents, vision-language environments (the shipped geo3k recipe), LLM-as-judge reward loops, and on-policy distillation onto a smaller student.

Three components, one loop

Molt composes Ray for placement and asynchronous queues, vLLM for rollout, and NVIDIA AutoModel with FSDP2 for training. None of the three is forked, so upstream improvements arrive as a container pin rather than a rebase.

The runtime is an agent pool, a set of vLLM engines behind a request router, and a single trainable policy actor. A streaming pool keeps prompt groups in flight so engines never drain while the actor trains. Partial rollout pauses the engines, broadcasts actor shards over NCCL directly to each engine, and resumes retained requests instead of discarding them.

The agent is an ordinary program

An RL run names one Python module that exports an AgentRunner. Everything else is ordinary code, including the reward.

Two forms are supported. With Env, the framework owns the LLM loop in a Gymnasium-aligned step(). With ChatAgent, the user owns the loop through a stock OpenAI or Anthropic SDK. Molt launches a loopback server that speaks both wire protocols, and every request decodes server-side into one token-exact accumulation. When a long-horizon agent compacts its context and rewrites the prefix, the server seals the current segment and opens a fresh one automatically.

Never train on a token you did not generate

Three correctness invariants organize the design. Token identity: sampled token ids define the trajectory, not a retokenized transcript. Policy-version semantics: trainable tokens keep their behavior-policy log-probabilities, and asynchronous use is corrected per token behind a sequence-level gate. Forward consistency: rollout and actor must agree on model semantics.

For mixture-of-experts policies, the last invariant matters most. The rollout and training routers select experts independently, and small numerical differences can flip top-k choices. Molt applies rollout routing replay, where vLLM returns its per-token expert ids and the training forward replays them.

Interactive explainer


Key Takeaways

  • Molt is an Apache-2.0 agentic RL framework in about 8.6K lines of RL code, roughly 7× smaller than verl.
  • Ray, vLLM, and NVIDIA AutoModel are composed, never forked, so upstream releases arrive as a container pin.
  • Agents are plain Python; a stock OpenAI or Anthropic SDK trains as-is through a token-exact loopback server.
  • Throughput is statistically comparable to a Megatron-based stack, with the MoE-mismatch caveat disclosed.
  • Scale is a flag: the same loop runs a dense 4B model and a 700B MoE at --fsdp.ep_size 256.

Check out the Paper and Repo here. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well.

Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.? Connect with us

YOU MAY ALSO LIKE

No cloud, no GPUs, no problem: Liquid AI’s new model LFM2.5-2.6B brings powerful AI agents to devices as small as a Raspberry Pi

OpenAI’s Ring-Shaped Smart Speaker Will Reportedly Cost Between $300 And $400


Asif Razzaq is the CEO of Marktechpost Media Inc.. As a visionary entrepreneur and engineer, Asif is committed to harnessing the potential of Artificial Intelligence for social good. His most recent endeavor is the launch of an Artificial Intelligence Media Platform, Marktechpost, which stands out for its in-depth coverage of machine learning and deep learning news that is both technically sound and easily understandable by a wide audience. The platform boasts of over 2 million monthly views, illustrating its popularity among audiences.

Credit: Source link

ShareTweetSendSharePin

Related Posts

No cloud, no GPUs, no problem: Liquid AI’s new model LFM2.5-2.6B brings powerful AI agents to devices as small as a Raspberry Pi
AI & Technology

No cloud, no GPUs, no problem: Liquid AI’s new model LFM2.5-2.6B brings powerful AI agents to devices as small as a Raspberry Pi

August 6, 2026
OpenAI’s Ring-Shaped Smart Speaker Will Reportedly Cost Between 0 And 0
AI & Technology

OpenAI’s Ring-Shaped Smart Speaker Will Reportedly Cost Between $300 And $400

August 6, 2026
AMD Buys Taalas to Put Hard-Wired AI Models in Its Accelerator Roadmap – Unite.AI
AI & Technology

AMD Buys Taalas to Put Hard-Wired AI Models in Its Accelerator Roadmap – Unite.AI

August 6, 2026
Cloudflare Introduces Kitesurf: An Agent-First Web Browser That Runs Entirely in V8 Isolates on Cloudflare Workers
AI & Technology

Cloudflare Introduces Kitesurf: An Agent-First Web Browser That Runs Entirely in V8 Isolates on Cloudflare Workers

August 6, 2026
Next Post
Live updates: US-Iran war news; Trump calls off Iran strikes subject to deal being ‘rapidly’ reached – CNN

Live updates: US-Iran war news; Trump calls off Iran strikes subject to deal being ‘rapidly’ reached - CNN

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
House Homeland Security Panel Calls Altman In Over OpenAI Breach – Unite.AI

House Homeland Security Panel Calls Altman In Over OpenAI Breach – Unite.AI

August 3, 2026
Venezuelans worry they won’t recover remains of loved ones still under rubble after earthquakes

Venezuelans worry they won’t recover remains of loved ones still under rubble after earthquakes

August 2, 2026
Belgium’s Nicolas Raskin says win against U.S. was ‘justice’ after red card controversy

Belgium’s Nicolas Raskin says win against U.S. was ‘justice’ after red card controversy

August 2, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!