• bitcoinBitcoin(BTC)$86,167.000.84%
  • ethereumEthereum(ETH)$2,755.680.86%
  • tetherTether(USDT)$1.000.01%
  • binancecoinBNB(BNB)$789.93-0.56%
  • rippleXRP(XRP)$1.586.33%
  • usd-coinUSDC(USDC)$1.000.01%
  • solanaSolana(SOL)$117.66-0.42%
  • tronTRON(TRX)$0.342957-0.53%
  • zcashZcash(ZEC)$1,514.52-1.38%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.01-0.60%
  • HyperliquidHyperliquid(HYPE)$95.380.61%
  • dogecoinDogecoin(DOGE)$0.1008378.00%
  • moneroMonero(XMR)$573.54-0.20%
  • whitebitWhiteBIT Coin(WBT)$86.690.80%
  • chainlinkChainlink(LINK)$13.080.30%
  • USDSUSDS(USDS)$1.000.01%
  • RainRain(RAIN)$0.013480-4.49%
  • cardanoCardano(ADA)$0.2530473.71%
  • leo-tokenLEO Token(LEO)$8.97-0.13%
  • stellarStellar(XLM)$0.2152582.65%
  • bitcoin-cashBitcoin Cash(BCH)$326.4221.31%
  • nearNEAR Protocol(NEAR)$4.457.72%
  • uniswapUniswap(UNI)$9.051.70%
  • Ethena USDeEthena USDe(USDE)$1.00-0.02%
  • avalanche-2Avalanche(AVAX)$11.01-2.29%
  • litecoinLitecoin(LTC)$62.10-0.63%
  • CantonCanton(CC)$0.1172271.84%
  • daiDai(DAI)$1.000.00%
  • USD1USD1(USD1)$1.00-0.03%
  • hedera-hashgraphHedera(HBAR)$0.0967706.20%
  • suiSui(SUI)$1.02-1.98%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.430.30%
  • BittensorBittensor(TAO)$320.3711.87%
  • shiba-inuShiba Inu(SHIB)$0.0000065.43%
  • crypto-com-chainCronos(CRO)$0.0668134.91%
  • Global DollarGlobal Dollar(USDG)$1.000.01%
  • MemeCoreMemeCore(M)$1.32-10.67%
  • paypal-usdPayPal USD(PYUSD)$1.000.01%
  • tether-goldTether Gold(XAUT)$4,334.57-0.34%
  • okbOKB(OKB)$122.32-0.12%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • Ripple USDRipple USD(RLUSD)$1.000.02%
  • BitwayBitway(BTW)$0.862.25%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.150.03%
  • aaveAave(AAVE)$144.29-1.11%
  • mantleMantle(MNT)$0.664.56%
  • EthenaEthena(ENA)$0.210783-6.06%
  • Pump.funPump.fun(PUMP)$0.0045213.35%
  • OndoOndo(ONDO)$0.433635-4.24%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

MiroMind-M1: Advancing Open-Source Mathematical Reasoning via Context-Aware Multi-Stage Reinforcement Learning

July 30, 2025
in AI & Technology
Reading Time: 5 mins read
A A
MiroMind-M1: Advancing Open-Source Mathematical Reasoning via Context-Aware Multi-Stage Reinforcement Learning
ShareShareShareShareShare

Large language models (LLMs) have recently demonstrated remarkable progress in multi-step reasoning, establishing mathematical problem-solving as a rigorous benchmark for assessing advanced capabilities. While proprietary models like GPT-4o and Claude Sonnet 4 lead performance, their closed-source nature impedes transparency and reproducibility. Addressing these gaps, MiroMind AI Released the MiroMind-M1 series, a fully open-source pipeline—spanning datasets, models, training code, and evaluation scripts—that sets new standards for openness and state-of-the-art mathematical reasoning within the Qwen-2.5 model ecosystem.

Architectural Foundation and Motivation

MiroMind-M1 is built on the robust Qwen-2.5 backbone, with enhancements geared explicitly for mathematical reasoning. The team adopts a two-stage training protocol:

YOU MAY ALSO LIKE

How To Enter VR Mode On Steam

Peloton Has Made A Foldable (Treadmill)

  1. Supervised Fine-Tuning (SFT): The model is fine-tuned on 719K carefully curated and verified mathematical problems, equipping it with strong step-by-step reasoning abilities.
  2. Reinforcement Learning with Verifiable Rewards (RLVR): Next, the model undergoes RL on 62K challenging and rigorously verifiable math problems, leveraging reward signals from a robust external verifier.

This approach is motivated by both the need for strong mathematical logic and by the lessons learned from leading RLMs: imitating chain-of-thought exemplars improves general reasoning, while reinforcement learning, guided by precise rewards, further refines accuracy and efficiency.

Data Transparency and Quality

A hallmark of the MiroMind-M1 project is the full openness and cleanliness of its training data:

  • SFT corpus composition: Draws from OpenR1, OpenThoughts, Light-R1, and Synthetic-1, ensuring problems have verified solutions and rich, multi-step reasoning traces.
  • Stringent deduplication and decontamination: Employs N-gram overlap filtering to eliminate duplication and data leakage with evaluation sets (e.g., AIME24, AIME25, MATH500).
  • Preference for long trajectories: Experiments show that training on samples with longer reasoning traces consistently yields higher benchmark scores, highlighting the importance of deep semantic content in the reasoning signal.

The resulting dataset provides 719K verified training traces—significantly advancing open reproducible research over prior efforts.

Supervised Fine-Tuning: Empirical Excellence

For SFT, MiroMind-SFT-7B is initialized from Qwen2.5-Math-7B and trained with a large context window (max 32,768 tokens) and a no-packing strategy to avoid cross-sample attention contamination. Its performance on key math benchmarks outpaces peer open models:

Model AIME24 AIME25 MATH500
DeepSeek-R1-Distill 55.5 40.4 92.8
MiMo-7B-SFT 58.7 44.3 93.0
MiroMind-SFT-7B 60.4 45.0 94.6

These results validate the efficacy of the data curation and training design: richer, deeper samples and no-packing lead to consistently superior performance.

CAMPO: Context-Aware Multi-Stage Policy Optimization

A key innovation in MiroMind-M1’s RLVR phase is the CAMPO algorithm. CAMPO addresses two critical RL challenges—training instability and token inefficiency—by:

  • Multi-stage training with expanding context limits: Training starts with constrained output lengths (e.g., 16K tokens), then gradually increases to allow deeper reasoning, balancing efficiency and thoroughness.
  • Dynamic repetition penalty: A dedicated repetition critic penalizes outputs exhibiting early or excessive repetition, preventing utility collapse and enforcing output diversity.
  • Accurate external verifier: The reward feedback system is substantially improved to robustly score math answers (including tricky cases with units, π, and percentages), ensuring training signals are tightly aligned with true correctness.

CAMPO not only stabilizes RL dynamics but also results in models that solve problems with fewer, more relevant tokens—accelerating inference and reducing costs without sacrificing accuracy.

Benchmark Performance: State-of-the-Art Efficiency

MiroMind’s open models achieve highly competitive or state-of-the-art results for open Qwen-2.5-based math models (7B/32B parameters):

Model AIME24 AIME25 MATH500
DeepSeek-R1-7B 55.5 39.2 –
MiMo-7B-RL 68.2 55.4 95.8
Skywork-OR1-7B 72.2 54.6 –
MiroMind-RL-7B 73.4 57.8 96.7
Skywork-OR1-32B 77.1 68.2 97.5
MiroMind-RL-32B 77.5 65.6 96.4

Notably, MiroMind-M1-RL models not only match or exceed peer accuracy, but do so with greater token efficiency—the 32B model produces shorter, more concise solutions without loss of correctness, thanks to CAMPO’s training.

Full Stack and Reproducibility

Every component of the MiroMind-M1 stack is openly released:

  • Model weights (SFT and RL checkpoints for both 7B and 32B scales)
  • Datasets (full 719K SFT, 62K RLVR)
  • Training scripts (supporting multi-node distributed training on Ray)
  • Evaluation code (standardized scripts and benchmark configs)

Researchers can replicate, audit, and extend MiroMind-M1 from raw data to trained models, advancing reproducibility and accelerating new open research.

Conclusion

MiroMind-M1 demonstrates that with careful data curation, innovative RL algorithms (CAMPO), and radical transparency, open-source language models can rival proprietary systems in advanced mathematical reasoning. This project sets a new bar for reproducibility and collaborative advancement in reasoning LLMs, providing both a high-quality resource and a robust platform for future innovation.


Check out the Paper, GitHub Page and Model on Hugging Face. All credit for this research goes to the researchers of this project. Also, feel free to follow us on Twitter and don’t forget to join our 100k+ ML SubReddit and Subscribe to our Newsletter.


Nikhil is an intern consultant at Marktechpost. He is pursuing an integrated dual degree in Materials at the Indian Institute of Technology, Kharagpur. Nikhil is an AI/ML enthusiast who is always researching applications in fields like biomaterials and biomedical science. With a strong background in Material Science, he is exploring new advancements and creating opportunities to contribute.

Credit: Source link

ShareTweetSendSharePin

Related Posts

How To Enter VR Mode On Steam
AI & Technology

How To Enter VR Mode On Steam

September 22, 2026
Peloton Has Made A Foldable (Treadmill)
AI & Technology

Peloton Has Made A Foldable (Treadmill)

September 22, 2026
OpenAI Faces Lawsuit From British Columbia Over Tumbler Ridge Shooting
AI & Technology

OpenAI Faces Lawsuit From British Columbia Over Tumbler Ridge Shooting

September 22, 2026
NVIDIA Introduces SoL-Pi: Auto-Research Loops That Cut Coding Agent Token Traffic by Up to 49%
AI & Technology

NVIDIA Introduces SoL-Pi: Auto-Research Loops That Cut Coding Agent Token Traffic by Up to 49%

September 22, 2026
Next Post
Beachgoers on high alert as shark sightings rise on East Coast

Beachgoers on high alert as shark sightings rise on East Coast

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Will Lindsay Clancy be retried? How and when it could happen

Will Lindsay Clancy be retried? How and when it could happen

September 17, 2026
Kornacki: Republicans hoping to make New Hampshire’s Senate race a ‘battleground’ in November

Kornacki: Republicans hoping to make New Hampshire’s Senate race a ‘battleground’ in November

September 15, 2026
Anthropic Taps Accenture’s Faculty for Embedded AI Model Evaluation – Unite.AI

Anthropic Taps Accenture’s Faculty for Embedded AI Model Evaluation – Unite.AI

September 18, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!