• bitcoinBitcoin(BTC)$83,339.00-1.31%
  • ethereumEthereum(ETH)$2,677.45-0.45%
  • tetherTether(USDT)$1.000.00%
  • binancecoinBNB(BNB)$762.93-1.48%
  • rippleXRP(XRP)$1.49-1.47%
  • usd-coinUSDC(USDC)$1.000.00%
  • solanaSolana(SOL)$118.67-2.51%
  • tronTRON(TRX)$0.3344890.24%
  • zcashZcash(ZEC)$1,530.74-3.20%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.060.00%
  • HyperliquidHyperliquid(HYPE)$88.24-3.46%
  • dogecoinDogecoin(DOGE)$0.093351-3.68%
  • chainlinkChainlink(LINK)$14.432.19%
  • moneroMonero(XMR)$532.15-1.78%
  • whitebitWhiteBIT Coin(WBT)$83.27-1.21%
  • USDSUSDS(USDS)$1.000.00%
  • cardanoCardano(ADA)$0.244396-3.91%
  • RainRain(RAIN)$0.0125410.03%
  • leo-tokenLEO Token(LEO)$9.060.43%
  • stellarStellar(XLM)$0.2203012.50%
  • nearNEAR Protocol(NEAR)$4.91-5.44%
  • bitcoin-cashBitcoin Cash(BCH)$309.51-6.56%
  • hedera-hashgraphHedera(HBAR)$0.12586035.02%
  • uniswapUniswap(UNI)$8.83-8.26%
  • litecoinLitecoin(LTC)$70.05-1.52%
  • CantonCanton(CC)$0.126908-5.23%
  • Ethena USDeEthena USDe(USDE)$1.000.01%
  • avalanche-2Avalanche(AVAX)$10.33-6.17%
  • suiSui(SUI)$1.15-7.11%
  • daiDai(DAI)$1.000.02%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.61-3.15%
  • USD1USD1(USD1)$1.000.00%
  • BittensorBittensor(TAO)$299.87-6.83%
  • crypto-com-chainCronos(CRO)$0.0676970.99%
  • shiba-inuShiba Inu(SHIB)$0.000006-4.27%
  • Global DollarGlobal Dollar(USDG)$1.000.04%
  • quant-networkQuant(QNT)$220.9223.82%
  • tether-goldTether Gold(XAUT)$4,127.23-3.53%
  • paypal-usdPayPal USD(PYUSD)$1.000.00%
  • MemeCoreMemeCore(M)$1.16-1.20%
  • EthenaEthena(ENA)$0.258756-5.51%
  • BitwayBitway(BTW)$0.96-17.55%
  • OndoOndo(ONDO)$0.51-5.46%
  • Ripple USDRipple USD(RLUSD)$1.000.02%
  • okbOKB(OKB)$117.49-3.16%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • Pump.funPump.fun(PUMP)$0.0051114.82%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.15-0.06%
  • aaveAave(AAVE)$145.94-4.92%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

A New MIT Study Shows Reinforcement Learning Minimizes Catastrophic Forgetting Compared to Supervised Fine-Tuning

September 8, 2025
in AI & Technology
Reading Time: 8 mins read
A A
A New MIT Study Shows Reinforcement Learning Minimizes Catastrophic Forgetting Compared to Supervised Fine-Tuning
ShareShareShareShareShare

What is catastrophic forgetting in foundation models?

Foundation models excel in diverse domains but are largely static once deployed. Fine-tuning on new tasks often introduces catastrophic forgetting—the loss of previously learned capabilities. This limitation poses a barrier for building long-lived, continually improving AI agents.

Why does online reinforcement learning forget less than supervised fine-tuning?

A new MIT study compares reinforcement learning (RL) and supervised fine-tuning (SFT). Both can achieve high performance on new tasks, but SFT tends to overwrite prior abilities. RL, by contrast, preserves them. The key lies in how each method shifts the model’s output distribution relative to the base policy.

YOU MAY ALSO LIKE

How To Improve Your Samsung Galaxy’s Battery Performance

A Modular, Repairable GPS Watch Is A Good First Step

https://arxiv.org/pdf/2509.04259

How can forgetting be measured?

The research team proposes an empirical forgetting law:

Forgetting∝KL(π0​∣∣π)

where π0 is the base model and π is the fine-tuned model. The forward KL divergence, measured on the new task, strongly predicts the extent of forgetting. This makes forgetting quantifiable without needing data from prior tasks.

What do experiments on large language models reveal?

Using Qwen 2.5 3B-Instruct as the base model, fine-tuning was performed on:

  • Math reasoning (Open-Reasoner-Zero),
  • Science Q&A (SciKnowEval subset),
  • Tool use (ToolAlpaca).

Performance was evaluated on prior benchmarks such as HellaSwag, MMLU, TruthfulQA, and HumanEval. Results showed that RL improved new-task accuracy while keeping prior-task accuracy stable, whereas SFT consistently sacrificed prior knowledge.

How does RL compare to SFT in robotics tasks?

In robotic control experiments with OpenVLA-7B fine-tuned in SimplerEnv pick-and-place scenarios, RL adaptation maintained general manipulation skills across tasks. SFT, while successful on the new task, degraded prior manipulation abilities—again illustrating RL’s conservatism in preserving knowledge.

What insights come from the ParityMNIST study?

To isolate mechanisms, the research team introduced a toy problem, ParityMNIST. Here, RL and SFT both reached high new-task accuracy, but SFT induced sharper declines on the FashionMNIST auxiliary benchmark. Crucially, plotting forgetting against KL divergence revealed a single predictive curve, validating KL as the governing factor.

Why do on-policy updates matter?

On-policy RL samples from the model’s own outputs, incrementally reweighting them by reward. This process constrains learning to distributions already close to the base model. SFT, in contrast, optimizes against fixed labels that may be arbitrarily distant. Theoretical analysis shows policy gradients converge to KL-minimal optimal solutions, formalizing RL’s advantage.

Are other explanations sufficient?

The research team tested alternatives: weight-space changes, hidden representation drift, sparsity of updates, and alternative distributional metrics (reverse KL, total variation, L2 distance). None matched the predictive strength of forward KL divergence, reinforcing that distributional closeness is the critical factor.

What are the broader implications?

  • Evaluation: Post-training should consider KL-conservatism, not just task accuracy.
  • Hybrid methods: Combining SFT efficiency with explicit KL minimization could yield optimal trade-offs.
  • Continual learning: RL’s Razor offers a measurable criterion for designing adaptive agents that learn new skills without erasing old ones.

Conclusion

The MIT research reframes catastrophic forgetting as a distributional problem governed by forward KL divergence. Reinforcement learning forgets less because its on-policy updates naturally bias toward KL-minimal solutions. This principle—RL’s Razor—provides both an explanation for RL’s robustness and a roadmap for developing post-training methods that support lifelong learning in foundation models.

Key Takeaways

  • Reinforcement learning (RL) preserves prior knowledge better than Supervised fine-tuning (SFT): Even when both achieve the same accuracy on new tasks, RL retains prior capabilities while SFT erases them.
  • Forgetting is predictable by KL divergence: The degree of catastrophic forgetting is strongly correlated with the forward KL divergence between the fine-tuned and base policy, measured on the new task.
  • RL’s Razor principle: On-policy RL converges to KL-minimal solutions, ensuring updates remain close to the base model and reducing forgetting.
  • Empirical validation across domains: Experiments on LLMs (math, science Q&A, tool use) and robotics tasks confirm RL’s robustness against forgetting, while SFT consistently trades old knowledge for new-task performance.
  • Controlled experiments confirm generality: In the ParityMNIST toy setting, both RL and SFT showed forgetting aligned with KL divergence, proving the principle holds beyond large-scale models.
  • Future design axis for post-training: Algorithms should be evaluated not only by new-task accuracy but also by how conservatively they shift distributions in KL space, opening avenues for hybrid RL–SFT methods.

Check out the PAPER and PROJECT PAGE. Feel free to check out our GitHub Page for Tutorials, Codes and Notebooks. Also, feel free to follow us on Twitter and don’t forget to join our 100k+ ML SubReddit and Subscribe to our Newsletter.



Michal Sutter is a data science professional with a Master of Science in Data Science from the University of Padova. With a solid foundation in statistical analysis, machine learning, and data engineering, Michal excels at transforming complex datasets into actionable insights.

Credit: Source link

ShareTweetSendSharePin

Related Posts

How To Improve Your Samsung Galaxy’s Battery Performance
AI & Technology

How To Improve Your Samsung Galaxy’s Battery Performance

September 28, 2026
A Modular, Repairable GPS Watch Is A Good First Step
AI & Technology

A Modular, Repairable GPS Watch Is A Good First Step

September 28, 2026
Fireworks AI Releases Ember-1: A Post-Trained Kimi K3 That Uses About 40% Fewer Tokens
AI & Technology

Fireworks AI Releases Ember-1: A Post-Trained Kimi K3 That Uses About 40% Fewer Tokens

September 28, 2026
You Can Now Preorder The Tiny Boox Picco Ereader
AI & Technology

You Can Now Preorder The Tiny Boox Picco Ereader

September 28, 2026
Next Post
Donald Trump drew boos and some cheers at the U.S. Open. Here’s what wasn’t shown on TV – The Athletic – The New York Times

Donald Trump drew boos and some cheers at the U.S. Open. Here’s what wasn’t shown on TV - The Athletic - The New York Times

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
LIVE: Day 20 of Lindsay Clancy murder trial | NBC News

LIVE: Day 20 of Lindsay Clancy murder trial | NBC News

September 24, 2026
Current with Christine Romans – Aug. 25 | NBC News NOW

Current with Christine Romans – Aug. 25 | NBC News NOW

September 24, 2026
Warrior Met Coal: A Volume Story Still Priced Like A Coal Price Bet

Warrior Met Coal: A Volume Story Still Priced Like A Coal Price Bet

September 27, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!