• bitcoinBitcoin(BTC)$86,507.000.68%
  • ethereumEthereum(ETH)$2,753.160.15%
  • tetherTether(USDT)$1.00-0.01%
  • binancecoinBNB(BNB)$787.61-1.10%
  • rippleXRP(XRP)$1.574.84%
  • usd-coinUSDC(USDC)$1.00-0.01%
  • solanaSolana(SOL)$117.920.28%
  • tronTRON(TRX)$0.341265-0.87%
  • zcashZcash(ZEC)$1,545.634.37%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.03-0.70%
  • HyperliquidHyperliquid(HYPE)$96.794.42%
  • dogecoinDogecoin(DOGE)$0.0996621.30%
  • moneroMonero(XMR)$572.73-0.47%
  • whitebitWhiteBIT Coin(WBT)$86.780.40%
  • chainlinkChainlink(LINK)$13.001.05%
  • USDSUSDS(USDS)$1.00-0.02%
  • RainRain(RAIN)$0.013194-6.27%
  • cardanoCardano(ADA)$0.2487532.19%
  • leo-tokenLEO Token(LEO)$8.981.10%
  • stellarStellar(XLM)$0.2141672.95%
  • bitcoin-cashBitcoin Cash(BCH)$329.1324.89%
  • nearNEAR Protocol(NEAR)$4.4611.45%
  • uniswapUniswap(UNI)$9.153.63%
  • Ethena USDeEthena USDe(USDE)$1.00-0.02%
  • avalanche-2Avalanche(AVAX)$11.030.46%
  • litecoinLitecoin(LTC)$61.80-0.47%
  • daiDai(DAI)$1.00-0.01%
  • CantonCanton(CC)$0.112976-1.79%
  • USD1USD1(USD1)$1.00-0.02%
  • hedera-hashgraphHedera(HBAR)$0.0959084.66%
  • suiSui(SUI)$1.01-0.19%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.440.54%
  • BittensorBittensor(TAO)$313.9510.53%
  • shiba-inuShiba Inu(SHIB)$0.0000061.24%
  • crypto-com-chainCronos(CRO)$0.0666254.43%
  • Global DollarGlobal Dollar(USDG)$1.00-0.01%
  • MemeCoreMemeCore(M)$1.32-10.88%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.02%
  • tether-goldTether Gold(XAUT)$4,338.93-0.04%
  • okbOKB(OKB)$122.670.38%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • Ripple USDRipple USD(RLUSD)$1.000.01%
  • BitwayBitway(BTW)$0.86-7.84%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.150.02%
  • aaveAave(AAVE)$144.070.71%
  • mantleMantle(MNT)$0.663.63%
  • OndoOndo(ONDO)$0.432363-2.67%
  • Pump.funPump.fun(PUMP)$0.0045064.95%
  • EthenaEthena(ENA)$0.204650-2.30%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

RA3: Mid-Training with Temporal Action Abstractions for Faster Reinforcement Learning (RL) Post-Training in Code LLMs

October 9, 2025
in AI & Technology
Reading Time: 4 mins read
A A
RA3: Mid-Training with Temporal Action Abstractions for Faster Reinforcement Learning (RL) Post-Training in Code LLMs
ShareShareShareShareShare

TL;DR: A new research from Apple, formalizes what “mid-training” should do before reinforcement learning RL post-training and introduces RA3 (Reasoning as Action Abstractions)—an EM-style procedure that learns temporally consistent latent actions from expert traces, then fine-tunes on those bootstrapped traces. It shows mid-training should (1) prune to a compact near-optimal action subspace and (2) shorten the effective planning horizon, improving RL convergence. Empirically, RA3 improves HumanEval/MBPP by ~8/4 points over base/NTP and accelerates RLVR on HumanEval+, MBPP+, LiveCodeBench, and Codeforces.

What does the research present?

The research team present the first formal treatment of how mid-training shapes post-training reinforcement learning RL: they breakdown outcomes into (i) pruning efficiency—how well mid-training selects a compact near-optimal action subset that shapes the initial policy prior—and (ii) RL convergence—how quickly post-training improves within that restricted set. The analysis argues mid-training is most effective when the decision space is compact and the effective horizon is short, favoring temporal abstractions over primitive next-token actions.

YOU MAY ALSO LIKE

Do USB Extenders Really Work And Are They Safe To Use?

How To Enter VR Mode On Steam

https://arxiv.org/pdf/2509.25810

Algorithm: RA3 in one pass

RA3 derives a sequential variational lower bound (a temporal ELBO) and optimizes it with an EM-like loop:

  • E-step (latent discovery): use RL to infer temporally consistent latent structures (abstractions) aligned to expert sequences.
  • M-step (model update): perform next-token prediction on the bootstrapped, latent-annotated traces to make those abstractions part of the model’s policy.

Results: code generation and RLVR

On Python code tasks, the research team reports that across multiple base models, RA3 improves average pass@k on HumanEval and MBPP by ~8 and ~4 points over the base model and an NTP mid-training baseline. In post-training, RLVR converges faster and to higher final performance on HumanEval+, MBPP+, LiveCodeBench, and Codeforces when initialized from RA3. These are mid- and post-training effects respectively; the evaluation scope is code generation.

Key Takeaways

  1. The research team formalizes mid-training via two determinants—pruning efficiency and impact on RL convergence—arguing effectiveness rises when the decision space is compact and the effective horizon is short.
  2. RA3 optimizes a sequential variational lower bound by iteratively discovering temporally consistent latent structures with RL and then fine-tuning on bootstrapped traces (EM-style).
  3. On code generation, RA3 reports ~+8 (HumanEval) and ~+4 (MBPP) average pass@k gains over base/NTP mid-training baselines across several model scales.
  4. Initializing post-training with RA3 accelerates RLVR convergence and improves asymptotic performance on HumanEval+, MBPP+, LiveCodeBench, and Codeforces.

Editorial Comments

RA3’s contribution is concrete and narrow: it formalizes mid-training around two determinants—pruning efficiency and RL convergence—and operationalizes them via a temporal ELBO optimized in an EM loop to learn persistent action abstractions before RLVR. The researchers report ~+8 (HumanEval) and ~+4 (MBPP) average pass@k gains over base/NTP and faster RLVR convergence on HumanEval+, MBPP+, LiveCodeBench, and Codeforces.


Check out the Technical Paper. Feel free to check out our GitHub Page for Tutorials, Codes and Notebooks. Also, feel free to follow us on Twitter and don’t forget to join our 100k+ ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well.

The post RA3: Mid-Training with Temporal Action Abstractions for Faster Reinforcement Learning (RL) Post-Training in Code LLMs appeared first on MarkTechPost.

Credit: Source link

ShareTweetSendSharePin

Related Posts

Do USB Extenders Really Work And Are They Safe To Use?
AI & Technology

Do USB Extenders Really Work And Are They Safe To Use?

September 22, 2026
How To Enter VR Mode On Steam
AI & Technology

How To Enter VR Mode On Steam

September 22, 2026
Peloton Has Made A Foldable (Treadmill)
AI & Technology

Peloton Has Made A Foldable (Treadmill)

September 22, 2026
OpenAI Faces Lawsuit From British Columbia Over Tumbler Ridge Shooting
AI & Technology

OpenAI Faces Lawsuit From British Columbia Over Tumbler Ridge Shooting

September 22, 2026
Next Post
Yankees left in disbelief: 'Didn't finish the goal' – ESPN

Yankees left in disbelief: 'Didn't finish the goal' - ESPN

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Berkshire Hathaway names Warren Buffett as chairman emeritus

Berkshire Hathaway names Warren Buffett as chairman emeritus

September 18, 2026
Lots of unsold units remain at One Wall Street — and pied-à-terre tax unlikely to help

Lots of unsold units remain at One Wall Street — and pied-à-terre tax unlikely to help

September 21, 2026
LIVE: Trump delivers remarks at the ‘Steel Across America’ event | NBC News

LIVE: Trump delivers remarks at the ‘Steel Across America’ event | NBC News

September 16, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!