• bitcoinBitcoin(BTC)$83,207.00-1.59%
  • ethereumEthereum(ETH)$2,652.84-2.03%
  • tetherTether(USDT)$1.000.00%
  • binancecoinBNB(BNB)$765.45-1.11%
  • rippleXRP(XRP)$1.49-2.41%
  • usd-coinUSDC(USDC)$1.000.00%
  • solanaSolana(SOL)$119.20-1.79%
  • tronTRON(TRX)$0.3335450.09%
  • zcashZcash(ZEC)$1,553.53-6.15%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.06-0.38%
  • HyperliquidHyperliquid(HYPE)$89.07-4.25%
  • dogecoinDogecoin(DOGE)$0.093540-3.37%
  • chainlinkChainlink(LINK)$13.90-2.32%
  • moneroMonero(XMR)$534.29-4.64%
  • whitebitWhiteBIT Coin(WBT)$83.01-1.64%
  • USDSUSDS(USDS)$1.00-0.01%
  • cardanoCardano(ADA)$0.247149-2.91%
  • RainRain(RAIN)$0.012551-1.29%
  • leo-tokenLEO Token(LEO)$9.080.65%
  • stellarStellar(XLM)$0.209982-2.92%
  • nearNEAR Protocol(NEAR)$5.17-3.91%
  • bitcoin-cashBitcoin Cash(BCH)$309.32-9.24%
  • uniswapUniswap(UNI)$9.23-7.75%
  • litecoinLitecoin(LTC)$70.87-1.99%
  • CantonCanton(CC)$0.1369991.41%
  • avalanche-2Avalanche(AVAX)$10.65-2.43%
  • suiSui(SUI)$1.211.97%
  • Ethena USDeEthena USDe(USDE)$1.000.02%
  • daiDai(DAI)$1.00-0.01%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.610.76%
  • USD1USD1(USD1)$1.000.01%
  • hedera-hashgraphHedera(HBAR)$0.0976513.93%
  • quant-networkQuant(QNT)$272.3553.88%
  • BitwayBitway(BTW)$1.3427.19%
  • BittensorBittensor(TAO)$306.95-5.56%
  • shiba-inuShiba Inu(SHIB)$0.000006-3.43%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • crypto-com-chainCronos(CRO)$0.064752-4.01%
  • tether-goldTether Gold(XAUT)$4,182.97-2.24%
  • OndoOndo(ONDO)$0.586.87%
  • paypal-usdPayPal USD(PYUSD)$1.000.00%
  • MemeCoreMemeCore(M)$1.18-3.91%
  • EthenaEthena(ENA)$0.265087-2.63%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • okbOKB(OKB)$117.37-3.09%
  • Pump.funPump.fun(PUMP)$0.00517816.63%
  • Circle USYCCircle USYC(USYC)$1.140.00%
  • aaveAave(AAVE)$150.23-4.03%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.14-0.14%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

ByteDance Introduces VAPO: A Novel Reinforcement Learning Framework for Advanced Reasoning Tasks

April 10, 2025
in AI & Technology
Reading Time: 4 mins read
A A
ByteDance Introduces VAPO: A Novel Reinforcement Learning Framework for Advanced Reasoning Tasks
ShareShareShareShareShare

In the Large Language Models (LLM) RL training, value-free methods like GRPO and DAPO have shown great effectiveness. The true potential lies in value-based methods, which allow more precise credit assignment by accurately tracing each action’s impact on subsequent returns. This precision is crucial for complex reasoning, where subtle errors can lead to catastrophic failures. However, training effective value models for long chain-of-thought (CoT) tasks face challenges: achieving low bias despite lengthy trajectories, managing distinct preferences of short and long responses, and addressing reward signal sparsity. Despite their theoretical advantages, these difficulties have hindered the full realization of value-based methods.

Value-based reinforcement learning methods for LLMs face three significant challenges when applied to long chain-of-thought reasoning tasks. First, the Value Model Bias issue identified in VC-PPO shows that initializing value models with reward models introduces positive bias. Second, Heterogeneous Sequence Lengths in complex reasoning tasks create difficulties for standard approaches like GAE with fixed parameters, which cannot effectively adapt to sequences ranging from very short to extremely long. Third, the Sparsity of the Reward Signal becomes problematic in verifier-based tasks that provide binary feedback rather than continuous values. This sparsity is worsened by lengthy CoT responses, creating a difficult exploration-exploitation trade-off during optimization.

YOU MAY ALSO LIKE

20 Agentic Use Cases of TypeSafe AI’s Jev

Google Research Introduces an AI Video Co-Director: 4 Agentic Frameworks for Coherent, Minutes-Long Video Generation

Researchers from ByteDance Seed have proposed Value Augmented Proximal Policy Optimization (VAPO), a value-based RL training framework to address the challenges of long CoT reasoning tasks. VAPO introduces three key innovations: a detailed value-based training framework with superior performance and efficiency, a Length-adaptive GAE mechanism that adjusts the parameter based on response lengths to optimize advantage estimation, and a systematic integration of techniques from prior research. VAPO combines these components to create a system where the collective improvements exceed what individual enhancements could achieve independently.  Using the Qwen2.5-32B model without SFT data, VAPO improves scores from 5 to 60, surpassing previous state-of-the-art methods by 10 points.

The VAPO is built upon the PPO algorithm with several key modifications to enhance mathematical reasoning capabilities. Training dynamics analysis reveals VAPO’s superior characteristics compared to DAPO, including smoother training curves indicating more stable optimization, better length scaling which enhances generalization capabilities, faster score growth due to the granular signals provided by the value model, and lower entropy in later training stages. While reduced entropy could potentially limit exploration, the method balances this trade-off effectively, resulting in minimal performance impact while improving reproducibility and stability. This shows how VAPO’s decisions directly address the core challenges of value-based RL in complex reasoning tasks.

While DeepSeek R1 using GRPO achieves 47 points on AIME24 and DAPO reaches 50 points, VAPO matches DAPO’s performance on Qwen-32b with just 60% of the update steps and achieves a new state-of-the-art score of 60.4 within only 5,000 steps. Vanilla PPO achieves only 5 points due to value model learning collapse, but VAPO finally achieves 60 points. Ablation studies validated the effectiveness of the seven proposed modifications: Value-Pretraining prevents collapse, decoupled GAE enables full optimization of long-form responses, adaptive GAE balances short and long response optimization, Clip-higher encourages thorough exploration, Token-level loss increases long response weighting, positive-example LM loss adds 6 points, and Group-Sampling contributes 5 points to the final performance.

In this paper, researchers introduced VAPO, an algorithm that utilizes the Qwen2.5-32B model to achieve state-of-the-art performance on the AIME24 benchmark. By introducing seven innovative techniques on top of the PPO framework, VAPO significantly refines value learning and creates an optimal balance between exploration and exploitation. This value-based approach decisively outperforms value-free methods like GRPO and DAPO, establishing a new performance ceiling for reasoning tasks. It addresses fundamental challenges in training value models for long CoT scenarios, providing a robust foundation for advancing LLMs in reasoning-intensive applications.


Check out the Paper. All credit for this research goes to the researchers of this project. Also, feel free to follow us on Twitter and don’t forget to join our 85k+ ML SubReddit.

🔥 [Register Now] miniCON Virtual Conference on OPEN SOURCE AI: FREE REGISTRATION + Certificate of Attendance + 3 Hour Short Event (April 12, 9 am- 12 pm PST) + Hands on Workshop [Sponsored]


Sana Hassan, a consulting intern at Marktechpost and dual-degree student at IIT Madras, is passionate about applying technology and AI to address real-world challenges. With a keen interest in solving practical problems, he brings a fresh perspective to the intersection of AI and real-life solutions.

Credit: Source link

ShareTweetSendSharePin

Related Posts

20 Agentic Use Cases of TypeSafe AI’s Jev
AI & Technology

20 Agentic Use Cases of TypeSafe AI’s Jev

September 28, 2026
Google Research Introduces an AI Video Co-Director: 4 Agentic Frameworks for Coherent, Minutes-Long Video Generation
AI & Technology

Google Research Introduces an AI Video Co-Director: 4 Agentic Frameworks for Coherent, Minutes-Long Video Generation

September 28, 2026
Which Is Better To Use?
AI & Technology

Which Is Better To Use?

September 28, 2026
Are 3D Printers Worth Buying In 2026?
AI & Technology

Are 3D Printers Worth Buying In 2026?

September 28, 2026
Next Post
An iconic sports car is getting an electric makeover — kind of

An iconic sports car is getting an electric makeover — kind of

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
FBI searches Eric Swalwell’s home and seizes his electronic devices

FBI searches Eric Swalwell’s home and seizes his electronic devices

September 26, 2026
Sen. Darline Graham thanks Trump after winning GOP Senate primary runoff in S.C.

Sen. Darline Graham thanks Trump after winning GOP Senate primary runoff in S.C.

September 23, 2026
London pays musical tribute to Dolly Parton

London pays musical tribute to Dolly Parton

September 22, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!