• bitcoinBitcoin(BTC)$81,017.004.61%
  • ethereumEthereum(ETH)$2,624.615.62%
  • tetherTether(USDT)$1.000.05%
  • binancecoinBNB(BNB)$761.471.19%
  • rippleXRP(XRP)$1.427.62%
  • usd-coinUSDC(USDC)$1.000.02%
  • solanaSolana(SOL)$112.566.57%
  • tronTRON(TRX)$0.3381690.72%
  • zcashZcash(ZEC)$1,535.311.26%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.030.41%
  • HyperliquidHyperliquid(HYPE)$92.415.51%
  • dogecoinDogecoin(DOGE)$0.0876573.94%
  • moneroMonero(XMR)$566.569.42%
  • whitebitWhiteBIT Coin(WBT)$83.154.29%
  • USDSUSDS(USDS)$1.000.02%
  • RainRain(RAIN)$0.0133915.48%
  • chainlinkChainlink(LINK)$12.395.38%
  • cardanoCardano(ADA)$0.2267506.17%
  • leo-tokenLEO Token(LEO)$8.89-0.10%
  • stellarStellar(XLM)$0.1946333.67%
  • uniswapUniswap(UNI)$8.975.65%
  • bitcoin-cashBitcoin Cash(BCH)$247.340.15%
  • nearNEAR Protocol(NEAR)$3.726.57%
  • Ethena USDeEthena USDe(USDE)$1.000.05%
  • daiDai(DAI)$1.00-0.01%
  • litecoinLitecoin(LTC)$58.266.20%
  • CantonCanton(CC)$0.1107261.71%
  • USD1USD1(USD1)$1.000.05%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.371.09%
  • avalanche-2Avalanche(AVAX)$8.477.02%
  • hedera-hashgraphHedera(HBAR)$0.0792713.08%
  • suiSui(SUI)$0.836.23%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • shiba-inuShiba Inu(SHIB)$0.0000051.88%
  • MemeCoreMemeCore(M)$1.300.89%
  • crypto-com-chainCronos(CRO)$0.0590250.41%
  • BittensorBittensor(TAO)$254.765.68%
  • paypal-usdPayPal USD(PYUSD)$1.000.03%
  • tether-goldTether Gold(XAUT)$4,374.500.21%
  • Circle USYCCircle USYC(USYC)$1.140.03%
  • okbOKB(OKB)$116.642.11%
  • Ripple USDRipple USD(RLUSD)$1.000.01%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.140.12%
  • aaveAave(AAVE)$143.006.88%
  • AsterAster(ASTER)$0.784.36%
  • mantleMantle(MNT)$0.615.55%
  • OndoOndo(ONDO)$0.4026254.90%
  • Pump.funPump.fun(PUMP)$0.004148-2.02%
  • MorphoMorpho(MORPHO)$2.7719.18%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

NVIDIA AI Releases ProRLv2: Advancing Reasoning in Language Models with Extended Reinforcement Learning RL

August 12, 2025
in AI & Technology
Reading Time: 6 mins read
A A
NVIDIA AI Releases ProRLv2: Advancing Reasoning in Language Models with Extended Reinforcement Learning RL
ShareShareShareShareShare

What Is ProRLv2?

ProRLv2 is the latest version of NVIDIA’s Prolonged Reinforcement Learning (ProRL), designed specifically to push the boundaries of reasoning in large language models (LLMs). By scaling reinforcement learning (RL) steps from 2,000 up to 3,000, ProRLv2 systematically tests how extended RL can unlock new solution spaces, creativity, and high-level reasoning that were previously inaccessible—even with smaller models like the 1.5B-parameter Nemotron-Research-Reasoning-Qwen-1.5B-v2.

Key Innovations in ProRLv2

ProRLv2 incorporates several innovations to overcome common RL limitations in LLM training:

  • REINFORCE++- Baseline: A robust RL algorithm that enables long-horizon optimization over thousands of steps, handling the instability typical in RL for LLMs.
  • KL Divergence Regularization & Reference Policy Reset: Periodically refreshes the reference model with the current best checkpoint, allowing stable progress and continued exploration by preventing the RL objective from dominating too early.
  • Decoupled Clipping & Dynamic Sampling (DAPO): Encourages diverse solution discovery by boosting unlikely tokens and focusing learning signals on prompts of intermediate difficulty.
  • Scheduled Length Penalty: Cyclically applied, helping maintain diversity and prevent entropy collapse as training lengthens.
  • Scaling Training Steps: ProRLv2 moves the RL training horizon from 2,000 to 3,000 steps, directly testing how much longer RL can expand reasoning abilities.

How ProRLv2 Expands LLM Reasoning

Nemotron-Research-Reasoning-Qwen-1.5B-v2, trained with ProRLv2 for 3,000 RL steps, sets a new standard for open-weight 1.5B models on reasoning tasks, including math, code, science, and logic puzzles:

  • Performance surpasses previous versions and competitors like DeepSeek-R1-1.5B.
  • Sustained gains with more RL steps: Longer training leads to continual improvements, especially on tasks where base models perform poorly, demonstrating genuine expansion in reasoning boundaries.
  • Generalization: Not only does ProRLv2 boost pass@1 accuracy, but it also enables novel reasoning and solution strategies on tasks not seen during training.
  • Benchmarks: Gains include average pass@1 improvements of 14.7% in math, 13.9% in coding, 54.8% in logic puzzles, 25.1% in STEM reasoning, and 18.1% in instruction-following tasks, with further improvements in v2 on unseen and harder benchmarks.

Why It Matters

The major finding of ProRLv2 is that continued RL training, with careful exploration and regularization, reliably expands what LLMs can learn and generalize. Rather than hitting an early plateau or overfitting, prolonged RL allows smaller models to rival much larger ones in reasoning—demonstrating that scaling RL itself is as important as model or dataset size.

YOU MAY ALSO LIKE

Consumers Sue Anthropic, OpenAI, SpaceXAI and Google Over Alleged AI Pact – Unite.AI

How Focus Mode Has Changed In iOS 27

Using Nemotron-Research-Reasoning-Qwen-1.5B-v2

The latest checkpoint is available for testing on Hugging Face. Loading the model:

from transformers import AutoTokenizer, AutoModelForCausalLM

tokenizer = AutoTokenizer.from_pretrained("nvidia/Nemotron-Research-Reasoning-Qwen-1.5B")
model = AutoModelForCausalLM.from_pretrained("nvidia/Nemotron-Research-Reasoning-Qwen-1.5B")

Conclusion

ProRLv2 redefines the limits of reasoning in language models by showing that RL scaling laws matter as much as size or data. Through advanced regularization and smart training schedules, it enables deep, creative, and generalizable reasoning even in compact architectures. The future lies in how far RL can push—not just how big models can get.


Check out the Unofficial Blog and Model on Hugging Face here. Feel free to check out our GitHub Page for Tutorials, Codes and Notebooks. Also, feel free to follow us on Twitter and don’t forget to join our 100k+ ML SubReddit and Subscribe to our Newsletter.


Nikhil is an intern consultant at Marktechpost. He is pursuing an integrated dual degree in Materials at the Indian Institute of Technology, Kharagpur. Nikhil is an AI/ML enthusiast who is always researching applications in fields like biomaterials and biomedical science. With a strong background in Material Science, he is exploring new advancements and creating opportunities to contribute.

Credit: Source link

ShareTweetSendSharePin

Related Posts

Consumers Sue Anthropic, OpenAI, SpaceXAI and Google Over Alleged AI Pact – Unite.AI
AI & Technology

Consumers Sue Anthropic, OpenAI, SpaceXAI and Google Over Alleged AI Pact – Unite.AI

September 19, 2026
How Focus Mode Has Changed In iOS 27
AI & Technology

How Focus Mode Has Changed In iOS 27

September 18, 2026
AI Almost Led The US Military To Start A War With China, Report Says
AI & Technology

AI Almost Led The US Military To Start A War With China, Report Says

September 18, 2026
Jina AI Releases jina-ocr-v1: A 3.4B MoE Document Parser With Built-In Speculative Decoding for Low-Budget GPUs
AI & Technology

Jina AI Releases jina-ocr-v1: A 3.4B MoE Document Parser With Built-In Speculative Decoding for Low-Budget GPUs

September 18, 2026
Next Post
Fritz the hippo celebrated his 3rd birthday with a watermelon cake at the Cincinnati Zoo

Fritz the hippo celebrated his 3rd birthday with a watermelon cake at the Cincinnati Zoo

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
9/11 survivors are being diagnosed with cancer 25 years after the attacks

9/11 survivors are being diagnosed with cancer 25 years after the attacks

September 13, 2026
Couple married in waist-deep water in the Philippines

Couple married in waist-deep water in the Philippines

September 17, 2026
Independent South Dakota candidate looks to ‘deny majority to either party’ in bid for Senate

Independent South Dakota candidate looks to ‘deny majority to either party’ in bid for Senate

September 17, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!