• bitcoinBitcoin(BTC)$84,497.000.40%
  • ethereumEthereum(ETH)$2,679.64-0.24%
  • tetherTether(USDT)$1.00-0.01%
  • binancecoinBNB(BNB)$777.460.72%
  • rippleXRP(XRP)$1.52-0.32%
  • usd-coinUSDC(USDC)$1.000.00%
  • solanaSolana(SOL)$122.550.96%
  • tronTRON(TRX)$0.333501-0.39%
  • zcashZcash(ZEC)$1,596.14-4.53%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.062.87%
  • HyperliquidHyperliquid(HYPE)$91.47-0.70%
  • dogecoinDogecoin(DOGE)$0.0968250.38%
  • chainlinkChainlink(LINK)$13.97-0.68%
  • moneroMonero(XMR)$546.66-1.54%
  • whitebitWhiteBIT Coin(WBT)$84.250.43%
  • USDSUSDS(USDS)$1.00-0.01%
  • cardanoCardano(ADA)$0.2548360.86%
  • RainRain(RAIN)$0.012562-2.56%
  • leo-tokenLEO Token(LEO)$9.050.98%
  • stellarStellar(XLM)$0.216208-0.41%
  • nearNEAR Protocol(NEAR)$5.5312.70%
  • bitcoin-cashBitcoin Cash(BCH)$332.54-1.03%
  • uniswapUniswap(UNI)$9.65-1.28%
  • litecoinLitecoin(LTC)$70.93-1.43%
  • CantonCanton(CC)$0.1366941.27%
  • suiSui(SUI)$1.279.76%
  • Ethena USDeEthena USDe(USDE)$1.000.01%
  • avalanche-2Avalanche(AVAX)$10.931.52%
  • daiDai(DAI)$1.000.00%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.622.62%
  • USD1USD1(USD1)$1.00-0.01%
  • hedera-hashgraphHedera(HBAR)$0.0945081.39%
  • BittensorBittensor(TAO)$326.592.73%
  • shiba-inuShiba Inu(SHIB)$0.000006-0.03%
  • quant-networkQuant(QNT)$230.9086.24%
  • crypto-com-chainCronos(CRO)$0.0670880.45%
  • BitwayBitway(BTW)$1.2218.72%
  • Global DollarGlobal Dollar(USDG)$1.00-0.01%
  • EthenaEthena(ENA)$0.2772722.43%
  • paypal-usdPayPal USD(PYUSD)$1.000.00%
  • MemeCoreMemeCore(M)$1.19-1.96%
  • OndoOndo(ONDO)$0.553.68%
  • tether-goldTether Gold(XAUT)$4,277.40-0.04%
  • okbOKB(OKB)$121.360.71%
  • Ripple USDRipple USD(RLUSD)$1.000.01%
  • Circle USYCCircle USYC(USYC)$1.140.00%
  • aaveAave(AAVE)$154.02-0.05%
  • Pump.funPump.fun(PUMP)$0.00507715.83%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.150.30%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Shanghai AI Lab Releases OREAL-7B and OREAL-32B: Advancing Mathematical Reasoning with Outcome Reward-Based Reinforcement Learning

February 11, 2025
in AI & Technology
Reading Time: 6 mins read
A A
Shanghai AI Lab Releases OREAL-7B and OREAL-32B: Advancing Mathematical Reasoning with Outcome Reward-Based Reinforcement Learning
ShareShareShareShareShare

Mathematical reasoning remains a difficult area for artificial intelligence (AI) due to the complexity of problem-solving and the need for structured, logical thinking. While large language models (LLMs) have made significant progress, they often struggle with tasks that require multi-step reasoning. Reinforcement learning (RL) has shown promise in improving these capabilities, yet traditional methods face challenges when rewards are sparse and binary, providing little feedback beyond a correct or incorrect answer.

Shanghai AI Laboratory has developed Outcome REwArd-based reinforcement Learning (OREAL), a series of mathematical reasoning models available as OREAL-7B and OREAL-32B. This framework is designed for situations where only binary rewards—correct or incorrect—are available. Unlike conventional RL approaches that rely on dense feedback, OREAL uses Best-of-N (BoN) sampling for behavior cloning and reshapes negative rewards to maintain gradient consistency.

YOU MAY ALSO LIKE

Bill Gates Says It’s ‘Completely Irresponsible’ For AI To Not Have Safeguards

Why The iPhone Duo Could Be Beneficial For Samsung’s Galaxy Z Fold 8

OREAL-7B and OREAL-32B demonstrate that smaller models can perform competitively with significantly larger models. OREAL-7B achieves a 94.0% pass@1 score on the MATH-500 benchmark, a result comparable to previous 32B models, while OREAL-32B reaches 95.0% pass@1, surpassing previous models trained through distillation.

Technical Insights and Advantages

The OREAL framework introduces several key techniques to improve mathematical reasoning:

  1. Best-of-N Sampling for Behavior Cloning: BoN sampling helps select optimal positive reasoning trajectories, allowing the model to learn from well-formed solutions.
  2. Reward Reshaping for Negative Samples: By adjusting negative rewards, the framework ensures gradient consistency between correct and incorrect samples, refining model optimization.
  3. Token-Level Reward Model for Chain-of-Thought Reasoning: Mathematical reasoning often involves long sequences of logical steps. OREAL assigns importance weights to key reasoning tokens, addressing the challenge of sparse binary feedback.
  4. On-Policy Reinforcement Learning: The model dynamically refines itself based on sampled queries, improving training efficiency and adaptability.

These techniques enable more stable training and better performance in long-sequence reasoning tasks, making reinforcement learning a viable alternative to traditional distillation approaches.

Performance and Evaluation

OREAL models have been tested across several benchmarks:

  • MATH-500 Benchmark:
    • OREAL-7B achieves 94.0% pass@1, a performance level previously seen only in 32B models.
    • OREAL-32B achieves 95.0% pass@1, setting a new standard in mathematical reasoning.
  • AIME2024 and OlympiadBench:
    • OREAL models outperform multiple baselines, showing strong generalization across problem types.
  • Comparison with OpenAI o-series and DeepSeek Models:
    • OREAL-32B surpasses DeepSeek-R1-Distill-Qwen-32B and OpenAI-o1-preview, demonstrating effective training strategies.
    • OREAL-7B achieves results on par with QwQ-32B-Preview and OpenAI-o1-mini, highlighting the impact of its reinforcement learning approach.

Conclusion

Shanghai AI Lab’s OREAL-7B and OREAL-32B models offer a refined approach to reinforcement learning in mathematical reasoning. By addressing the challenge of sparse binary rewards through Best-of-N sampling, reward shaping, and token-level importance weighting, these models achieve competitive performance even at smaller scales. The OREAL framework provides valuable insights into how reinforcement learning can be optimized for complex reasoning tasks, suggesting new directions for improving AI’s problem-solving capabilities in structured domains.


Check out the Paper, OREAL-7B and OREAL-32B. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter and join our Telegram Channel and LinkedIn Group. Don’t Forget to join our 75k+ ML SubReddit.

🚨 Recommended Open-Source AI Platform: ‘IntellAgent is a An Open-Source Multi-Agent Framework to Evaluate Complex Conversational AI System’ (Promoted)


Asif Razzaq is the CEO of Marktechpost Media Inc.. As a visionary entrepreneur and engineer, Asif is committed to harnessing the potential of Artificial Intelligence for social good. His most recent endeavor is the launch of an Artificial Intelligence Media Platform, Marktechpost, which stands out for its in-depth coverage of machine learning and deep learning news that is both technically sound and easily understandable by a wide audience. The platform boasts of over 2 million monthly views, illustrating its popularity among audiences.

✅ [Recommended] Join Our Telegram Channel

Credit: Source link

ShareTweetSendSharePin

Related Posts

Bill Gates Says It’s ‘Completely Irresponsible’ For AI To Not Have Safeguards
AI & Technology

Bill Gates Says It’s ‘Completely Irresponsible’ For AI To Not Have Safeguards

September 27, 2026
Why The iPhone Duo Could Be Beneficial For Samsung’s Galaxy Z Fold 8
AI & Technology

Why The iPhone Duo Could Be Beneficial For Samsung’s Galaxy Z Fold 8

September 27, 2026
How To Improve Your Router’s Security In 10 Minutes
AI & Technology

How To Improve Your Router’s Security In 10 Minutes

September 27, 2026
Humanoid Robots Are Getting Even Creepier (This One Can Cry On Command)
AI & Technology

Humanoid Robots Are Getting Even Creepier (This One Can Cry On Command)

September 27, 2026
Next Post
Adapting To Change: Getty Realty Stock And The EV Transition (NYSE:GTY)

Adapting To Change: Getty Realty Stock And The EV Transition (NYSE:GTY)

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Titanic-themed slide for kids makes headlines

Titanic-themed slide for kids makes headlines

September 27, 2026
Trump-Xi Summit Puts Global AI Race in Focus

Trump-Xi Summit Puts Global AI Race in Focus

September 24, 2026
A conversation with Marine veteran and advocate Jason Bush

A conversation with Marine veteran and advocate Jason Bush

September 25, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!