• bitcoinBitcoin(BTC)$77,494.00-1.31%
  • ethereumEthereum(ETH)$2,422.14-1.75%
  • tetherTether(USDT)$1.00-0.02%
  • binancecoinBNB(BNB)$683.04-1.14%
  • rippleXRP(XRP)$1.35-1.91%
  • usd-coinUSDC(USDC)$1.00-0.01%
  • solanaSolana(SOL)$100.09-2.87%
  • tronTRON(TRX)$0.322408-3.00%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.021.73%
  • HyperliquidHyperliquid(HYPE)$82.77-1.68%
  • zcashZcash(ZEC)$830.50-1.94%
  • dogecoinDogecoin(DOGE)$0.081758-1.28%
  • RainRain(RAIN)$0.0168240.74%
  • USDSUSDS(USDS)$1.00-0.01%
  • moneroMonero(XMR)$501.17-2.67%
  • leo-tokenLEO Token(LEO)$9.37-1.99%
  • whitebitWhiteBIT Coin(WBT)$71.30-1.48%
  • chainlinkChainlink(LINK)$11.23-0.73%
  • cardanoCardano(ADA)$0.196531-0.69%
  • stellarStellar(XLM)$0.175825-0.73%
  • bitcoin-cashBitcoin Cash(BCH)$246.00-0.18%
  • daiDai(DAI)$1.000.00%
  • CantonCanton(CC)$0.113815-6.40%
  • Ethena USDeEthena USDe(USDE)$1.00-0.03%
  • USD1USD1(USD1)$1.00-0.02%
  • litecoinLitecoin(LTC)$49.592.29%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.32-4.92%
  • uniswapUniswap(UNI)$5.8512.69%
  • Global DollarGlobal Dollar(USDG)$1.00-0.03%
  • hedera-hashgraphHedera(HBAR)$0.0740630.65%
  • avalanche-2Avalanche(AVAX)$7.22-0.07%
  • shiba-inuShiba Inu(SHIB)$0.0000051.55%
  • suiSui(SUI)$0.72-0.66%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.03%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • crypto-com-chainCronos(CRO)$0.054862-3.00%
  • tether-goldTether Gold(XAUT)$4,336.63-2.23%
  • nearNEAR Protocol(NEAR)$1.90-1.06%
  • MemeCoreMemeCore(M)$1.06-2.70%
  • Ripple USDRipple USD(RLUSD)$1.00-0.01%
  • okbOKB(OKB)$110.40-1.40%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.14-0.11%
  • BittensorBittensor(TAO)$219.81-4.65%
  • aaveAave(AAVE)$127.682.97%
  • pax-goldPAX Gold(PAXG)$4,343.67-2.23%
  • AsterAster(ASTER)$0.69-1.13%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.057005-0.83%
  • MorphoMorpho(MORPHO)$2.603.33%
  • mantleMantle(MNT)$0.53-2.40%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Sigmoidal Scaling Curves Make Reinforcement Learning RL Post-Training Predictable for LLMs

October 18, 2025
in AI & Technology
Reading Time: 7 mins read
A A
Sigmoidal Scaling Curves Make Reinforcement Learning RL Post-Training Predictable for LLMs
ShareShareShareShareShare

Reinforcement Learning RL post-training is now a major lever for reasoning-centric LLMs, but unlike pre-training, it hasn’t had predictive scaling rules. Teams pour tens of thousands of GPU-hours into runs without a principled way to estimate whether a recipe will keep improving with more compute. A new research from Meta, UT Austin, UCL, Berkeley, Harvard, and Periodic Labs provides a compute-performance framework—validated over >400,000 GPU-hours—that models RL progress with a sigmoidal curve and supplies a tested recipe, ScaleRL, that follows those predicted curves up to 100,000 GPU-hours.

Fit a sigmoid, not a power law

Pre-training often fits power laws (loss vs compute). RL fine-tuning targets bounded metrics (e.g., pass rate/mean reward). The research team show sigmoidal fits to pass rate vs training compute are empirically more robust and stable than power-law fits, especially when you want to extrapolate from smaller runs to larger budgets. They exclude the very early, noisy regime (~first 1.5k GPU-hours) and fit the predictable portion that follows. The sigmoidal parameters have intuitive roles: one sets the asymptotic performance (ceiling), another the efficiency/exponent, and another the midpoint where gains are fastest.

YOU MAY ALSO LIKE

Anthropic Releases Claude Fable 5.1 and Claude Mythos 5.1: 52.6% on Terminal-Bench-Science and 75% Cheaper Cache Reads

Frontier models can recover up to 65% of facts they can’t directly recall — just by thinking longer

https://arxiv.org/pdf/2510.13786

Why that matters: After ~1–2k GPU-hours, you can fit the curve and forecast whether pushing to 10k–100k GPU-hours is worth it—before you burn the budget. The research also shows power-law fits can produce misleading ceilings unless you only fit at very high compute, which defeats the purpose of early forecasting.

ScaleRL: a recipe that scales predictably

ScaleRL is not just new algorithm; it’s a composition of choices that produced stable, extrapolatable scaling in the study:

  • Asynchronous Pipeline RL (generator–trainer split across GPUs) for off-policy throughput.
  • CISPO (truncated importance-sampling REINFORCE) as the RL loss.
  • FP32 precision at the logits to avoid numeric mismatch between generator and trainer.
  • Prompt-level loss averaging and batch-level advantage normalization.
  • Forced length interruptions to cap runaway traces.
  • Zero-variance filtering (drop prompts that provide no gradient signal).
  • No-Positive-Resampling (remove high-pass-rate prompts ≥0.9 from later epochs).

The research team validated each component with leave-one-out (LOO) ablations at 16k GPU-hours and show that ScaleRL’s fitted curves reliably extrapolate from 8k → 16k, then hold at much larger scales—including a single run extended to 100k GPU-hours.

https://arxiv.org/pdf/2510.13786

Results and generalization

Two key demonstrations:

  1. Predictability at scale: For an 8B dense model and a Llama-4 17B×16 MoE (“Scout”), the extended training closely followed the sigmoid extrapolations derived from smaller-compute segments.
  2. Downstream transfer: Pass-rate improvements on an iid validation set track downstream evaluation (e.g., AIME-24), suggesting the compute-performance curve isn’t a dataset artifact.

The research also compares fitted curves for prevalent recipes (e.g., DeepSeek (GRPO), Qwen-2.5 (DAPO), Magistral, MiniMax-M1) and reports higher asymptotic performance and better compute efficiency for ScaleRL in their setup.

https://arxiv.org/pdf/2510.13786

Which knobs move the ceiling vs the efficiency?

The framework lets you classify design choices:

  • Ceiling movers (asymptote): scaling model size (e.g., MoE) and longer generation lengths (up to 32,768 tokens) raise the asymptotic performance but may slow early progress. Larger global batch size can also lift the final asymptote and stabilize training.
  • Efficiency shapers: loss aggregation, advantage normalization, data curriculum, and the off-policy pipeline mainly change how fast you approach the ceiling, not the ceiling itself.

Operationally, the research team advises fitting curves early and prioritizing interventions that raise the ceiling, then tune the efficiency knobs to reach it faster at fixed compute.

Key Takeaways

  • The research team models RL post-training progress with sigmoidal compute-performance curves (pass-rate vs. log compute), enabling reliable extrapolation—unlike power-law fits on bounded metrics.
  • A best-practice recipe, ScaleRL, combines PipelineRL-k (asynchronous generator–trainer), CISPO loss, FP32 logits, prompt-level aggregation, advantage normalization, interruption-based length control, zero-variance filtering, and no-positive-resampling.
  • Using these fits, the research team predicted and matched extended runs up to 100k GPU-hours (8B dense) and ~50k GPU-hours (17B×16 MoE “Scout”) on validation curves.
  • Ablations show some choices move the asymptotic ceiling (A) (e.g., model scale, longer generation lengths, larger global batch), while others mainly improve compute efficiency (B) (e.g., aggregation/normalization, curriculum, off-policy pipeline).
  • The framework provides early forecasting to decide whether to scale a run, and improvements on the in-distribution validation track downstream metrics (e.g., AIME-24), supporting external validity.

Editorial Comments

This work turns RL post-training from trial-and-error into forecastable engineering. It fits sigmoidal compute-performance curves (pass-rate vs. log compute) to predict returns and decide when to stop or scale. It also provides a concrete recipe, ScaleRL, that uses PipelineRL-style asynchronous generation/training, the CISPO loss, and FP32 logits for stability. The study reports >400,000 GPU-hours of experiments and a single-run extension to 100,000 GPU-hours. Results support a clean split: some choices raise the asymptote; others mainly improve compute efficiency. That separation helps teams prioritize ceiling-moving changes before tuning throughput knobs.


Check out the PAPER. Feel free to check out our GitHub Page for Tutorials, Codes and Notebooks. Also, feel free to follow us on Twitter and don’t forget to join our 100k+ ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well.

The post Sigmoidal Scaling Curves Make Reinforcement Learning RL Post-Training Predictable for LLMs appeared first on MarkTechPost.

Credit: Source link

ShareTweetSendSharePin

Related Posts

Anthropic Releases Claude Fable 5.1 and Claude Mythos 5.1: 52.6% on Terminal-Bench-Science and 75% Cheaper Cache Reads
AI & Technology

Anthropic Releases Claude Fable 5.1 and Claude Mythos 5.1: 52.6% on Terminal-Bench-Science and 75% Cheaper Cache Reads

September 1, 2026
Frontier models can recover up to 65% of facts they can’t directly recall — just by thinking longer
AI & Technology

Frontier models can recover up to 65% of facts they can’t directly recall — just by thinking longer

September 1, 2026
The New Street Fighter Movie Trailer Looks Fun In All The Right Ways
AI & Technology

The New Street Fighter Movie Trailer Looks Fun In All The Right Ways

September 1, 2026
Anthropic Announces Enterprise Frontier Safeguards, Customer-Held Data – Unite.AI
AI & Technology

Anthropic Announces Enterprise Frontier Safeguards, Customer-Held Data – Unite.AI

September 1, 2026
Next Post
Carson Beck throws 4 INTs as Louisville stuns No. 2 Miami – ESPN

Carson Beck throws 4 INTs as Louisville stuns No. 2 Miami - ESPN

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Meta to Pay Up to .1 Billion in Landmark Settlement Over Social Media Addiction Claims – The New York Times

Meta to Pay Up to $17.1 Billion in Landmark Settlement Over Social Media Addiction Claims – The New York Times

August 26, 2026
More Americans moving to the Midwest amid nationwide affordability crisis

More Americans moving to the Midwest amid nationwide affordability crisis

September 1, 2026
California manufacturing giant moves dozens of jobs to Nevada

California manufacturing giant moves dozens of jobs to Nevada

August 27, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!