• bitcoinBitcoin(BTC)$83,219.00-0.11%
  • ethereumEthereum(ETH)$2,667.080.60%
  • tetherTether(USDT)$1.00-0.01%
  • binancecoinBNB(BNB)$757.00-1.52%
  • rippleXRP(XRP)$1.49-0.05%
  • usd-coinUSDC(USDC)$1.000.00%
  • solanaSolana(SOL)$117.76-1.44%
  • tronTRON(TRX)$0.3341150.07%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.000.00%
  • zcashZcash(ZEC)$1,368.15-12.10%
  • HyperliquidHyperliquid(HYPE)$87.26-3.02%
  • dogecoinDogecoin(DOGE)$0.093253-0.74%
  • chainlinkChainlink(LINK)$14.826.97%
  • moneroMonero(XMR)$541.050.97%
  • whitebitWhiteBIT Coin(WBT)$83.170.13%
  • USDSUSDS(USDS)$1.00-0.04%
  • cardanoCardano(ADA)$0.243189-1.70%
  • RainRain(RAIN)$0.012416-1.11%
  • leo-tokenLEO Token(LEO)$9.01-0.28%
  • stellarStellar(XLM)$0.2254857.28%
  • bitcoin-cashBitcoin Cash(BCH)$306.23-3.02%
  • nearNEAR Protocol(NEAR)$4.66-10.69%
  • CantonCanton(CC)$0.134152-3.24%
  • uniswapUniswap(UNI)$8.58-7.62%
  • litecoinLitecoin(LTC)$67.84-4.27%
  • hedera-hashgraphHedera(HBAR)$0.11879724.34%
  • Ethena USDeEthena USDe(USDE)$1.00-0.01%
  • avalanche-2Avalanche(AVAX)$10.49-1.29%
  • daiDai(DAI)$1.000.00%
  • suiSui(SUI)$1.12-7.80%
  • USD1USD1(USD1)$1.00-0.02%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.56-3.36%
  • quant-networkQuant(QNT)$249.22-2.37%
  • BittensorBittensor(TAO)$301.80-1.75%
  • crypto-com-chainCronos(CRO)$0.0686366.10%
  • tether-goldTether Gold(XAUT)$4,133.64-1.52%
  • shiba-inuShiba Inu(SHIB)$0.000006-2.32%
  • Global DollarGlobal Dollar(USDG)$1.000.02%
  • BitwayBitway(BTW)$1.16-10.28%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.02%
  • EthenaEthena(ENA)$0.250172-6.59%
  • Ripple USDRipple USD(RLUSD)$1.00-0.01%
  • OndoOndo(ONDO)$0.51-11.72%
  • okbOKB(OKB)$118.550.49%
  • MemeCoreMemeCore(M)$1.08-8.76%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • aaveAave(AAVE)$149.32-0.40%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.150.18%
  • Pump.funPump.fun(PUMP)$0.004876-5.98%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Nebius AI Advances Open-Weight LLMs Through Reinforcement Learning for Capable SWE Agents

August 13, 2025
in AI & Technology
Reading Time: 9 mins read
A A
Nebius AI Advances Open-Weight LLMs Through Reinforcement Learning for Capable SWE Agents
ShareShareShareShareShare

The landscape of software engineering automation is evolving rapidly, driven by advances in Large Language Models (LLMs). However, most approaches to training capable agents rely on proprietary models or costly teacher-based methods, leaving open-weight LLMs with limited capabilities in real-world scenarios. A team of researchers from Nebius AI and Humanoid introduced a reinforcement learning framework for training long-context, multi-turn software engineering agents using a modified Decoupled Advantage Policy Optimization (DAPO) algorithm. The research explains a technical breakthrough in applying reinforcement learning (RL) to open-source LLMs for genuine, multi-turn software engineering tasks—moving beyond the single-turn, bandit-style settings that dominate RL for LLMs today.

Beyond Single-Turn Reinforcement Learning RL

Most RL methods for LLMs optimize for tasks such as mathematical reasoning or one-shot code generation, where agent actions are rewarded only at the conclusion and environments do not provide intermediate feedback. However, software engineering (SWE) is fundamentally different: it requires agents to operate over long sequences of actions, interpret rich feedback (compiler errors, test logs), and maintain context over hundreds of thousands of tokens—far exceeding typical single-step interaction loops.

Core Challenges in RL for SWE

  • Long-Horizon Reasoning: Agents must sustain logical coherence across many steps, often requiring context windows beyond 100k tokens.
  • Stateful Environment Feedback: Actions yield meaningful, non-trivial observations (e.g., shell command outputs, test suite results) that guide subsequent decisions.
  • Sparse/Delayed Rewards: Success signals typically emerge only at the end of complex interactions, complicating credit assignment.
  • Evaluation Complexity: Measuring progress requires full trajectory unrolling and can be noisy due to test flakiness.

The Technical Recipe: Modified DAPO and Agent Design

The research team demonstrates a two-stage learning pipeline for training a Qwen2.5-72B-Instruct agent:

1. Rejection Fine-Tuning (RFT)

The journey begins with supervised fine-tuning. The agent is run across 7,249 rigorously filtered SWE tasks (from the SWE-REBENCH dataset). Successful interaction traces—where the agent passes the environmental test suite—are used to fine-tune the model, particularly masking invalid environment-formatting actions during training. This alone boosts baseline accuracy from 11% to 20% on the SWE-bench Verified benchmark.

2. Reinforcement Learning Using Modified DAPO

Building on Decoupled Advantage Policy Optimization (DAPO), several key modifications are introduced for scalability and stability:

  • Asymmetric Clipping: Prevents collapse in policy entropy, maintaining exploration.
  • Dynamic Sample Filtering: Focuses optimization on trajectories with actual learning signal.
  • Length Penalties: Discourages excessive episode length, helping the agent avoid getting stuck in loops.
  • Token-Level Averaging: Every token in every trajectory contributes equally to the gradient, empowering longer trajectories to influence updates.

The agent utilizes a ReAct-style loop, which lets it combine reasoning steps with tool usage. Its supported toolkit includes arbitrary shell commands, precise code edits, navigation/search utilities, and a submit action to signal episode completion. Each interaction is grounded in a robust sandboxed environment, initialized from real repository snapshots and backed by a GitHub-style issue prompt.

Scaling to Long Contexts and Real Benchmarks

Initially trained with a context length of 65k tokens (already double that of most open models), performance stalls at 32%. A second RL phase expands the context to 131k tokens and doubles the episode length ceiling, focusing subsequent training on only the most beneficial tasks from the pool. This enables scaling to longer stack traces and diff histories inherent to real-world debugging and patching tasks.

Results: Closing the Gap with Baselines

  • The final RL-trained agent attains 39% Pass@1 accuracy on the SWE-bench Verified benchmark, doubling the rejection fine-tuned baseline, and matching the performance of cutting-edge open-weight models such as DeepSeek-V3-0324, all without teacher-based supervision.
  • On held-out SWE-rebench splits, scores remain competitive (35% for May, 31.7% for June), indicating the method’s robustness.
  • When compared head-to-head with top open baselines and specialized SWE agents, the RL agent matches or outperforms several models, confirming the effectiveness of the RL methodology in this domain.
Pass@1 SWE-bench Verified Pass@10 Pass@1 SWE-rebench May Pass@10
Qwen2.5-72B-Instruct (RL, final) 39.04% 58.4% 35.0% 52.5%
DeepSeek-V3-0324 39.56% 62.2% 36.75% 60.0%
Qwen3-235B no-thinking 25.84% 54.4% 27.25% 57.5%
Llama4 Maverick 15.84% 47.2% 19.0% 50.0%

Pass@1 scores are averaged over 10 runs and reported as mean ± standard error.

Key Insights

  • Credit Assignment: RL in this sparse-reward regime remains fundamentally challenging. The paper suggests future work with reward shaping, step-level critics, or prefix-based rollouts for more granular feedback.
  • Uncertainty Estimation: Real-world agents need to know when to abstain or express confidence. Techniques like output entropy or explicit confidence scoring are next steps.
  • Infrastructure: Training utilized context parallelism (splitting long sequences over GPUs) on 16 H200 nodes, with distributed orchestration via Kubernetes and Tracto AI, and vLLM for fast inference.

Conclusion

This research validates RL as a potent paradigm for building autonomous software engineers using open-weight LLMs. By conquering long-horizon, multi-turn, real-environment tasks, the methodology paves the way for scalable, teacher-free agent development—directly leveraging the power of interaction rather than static instruction. With further refinements, such RL pipelines promise efficient, reliable, and versatile automation for the future of software engineering.

YOU MAY ALSO LIKE

How To Get Started With Shortcuts On Your MacBook

The Warning Signs That Your iPhone Battery Needs To Be Replaced


Check out the Paper here. Feel free to check out our GitHub Page for Tutorials, Codes and Notebooks. Also, feel free to follow us on Twitter and don’t forget to join our 100k+ ML SubReddit and Subscribe to our Newsletter.


Nikhil is an intern consultant at Marktechpost. He is pursuing an integrated dual degree in Materials at the Indian Institute of Technology, Kharagpur. Nikhil is an AI/ML enthusiast who is always researching applications in fields like biomaterials and biomedical science. With a strong background in Material Science, he is exploring new advancements and creating opportunities to contribute.

Credit: Source link

ShareTweetSendSharePin

Related Posts

How To Get Started With Shortcuts On Your MacBook
AI & Technology

How To Get Started With Shortcuts On Your MacBook

September 29, 2026
The Warning Signs That Your iPhone Battery Needs To Be Replaced
AI & Technology

The Warning Signs That Your iPhone Battery Needs To Be Replaced

September 28, 2026
How To Improve Your Android Phone’s Battery Life
AI & Technology

How To Improve Your Android Phone’s Battery Life

September 28, 2026
Discord Is Testing A Lightweight Mode To Free Up Resources While Gaming
AI & Technology

Discord Is Testing A Lightweight Mode To Free Up Resources While Gaming

September 28, 2026
Next Post
Sky Harbour Group Corporation (SKYH) Q2 2025 Earnings Call Transcript

Sky Harbour Group Corporation (SKYH) Q2 2025 Earnings Call Transcript

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Americans missing in deadly flood disaster

Americans missing in deadly flood disaster

September 22, 2026
China, US agree to  billion tariff cut, AI dialogue during Xi visit, Beijing says – Reuters

China, US agree to $30 billion tariff cut, AI dialogue during Xi visit, Beijing says – Reuters

September 26, 2026
Family First credit union CEO out after ‘Lake America’ AI photo

Family First credit union CEO out after ‘Lake America’ AI photo

September 25, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!