• bitcoinBitcoin(BTC)$83,927.001.16%
  • ethereumEthereum(ETH)$2,714.432.09%
  • tetherTether(USDT)$1.000.00%
  • binancecoinBNB(BNB)$763.420.13%
  • rippleXRP(XRP)$1.511.29%
  • usd-coinUSDC(USDC)$1.000.01%
  • solanaSolana(SOL)$119.290.64%
  • tronTRON(TRX)$0.3353860.28%
  • zcashZcash(ZEC)$1,416.83-9.16%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.000.00%
  • HyperliquidHyperliquid(HYPE)$88.34-1.53%
  • dogecoinDogecoin(DOGE)$0.0949702.01%
  • chainlinkChainlink(LINK)$15.319.61%
  • moneroMonero(XMR)$541.232.20%
  • whitebitWhiteBIT Coin(WBT)$83.971.27%
  • USDSUSDS(USDS)$1.000.00%
  • cardanoCardano(ADA)$0.2508252.15%
  • RainRain(RAIN)$0.012473-1.07%
  • leo-tokenLEO Token(LEO)$9.060.74%
  • stellarStellar(XLM)$0.2295688.26%
  • bitcoin-cashBitcoin Cash(BCH)$311.450.55%
  • nearNEAR Protocol(NEAR)$4.77-5.89%
  • uniswapUniswap(UNI)$9.030.95%
  • litecoinLitecoin(LTC)$68.61-3.90%
  • CantonCanton(CC)$0.131051-1.56%
  • hedera-hashgraphHedera(HBAR)$0.1176871.45%
  • avalanche-2Avalanche(AVAX)$11.519.41%
  • Ethena USDeEthena USDe(USDE)$1.00-0.01%
  • suiSui(SUI)$1.17-0.50%
  • daiDai(DAI)$1.00-0.02%
  • USD1USD1(USD1)$1.000.00%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.57-6.10%
  • quant-networkQuant(QNT)$252.919.82%
  • BittensorBittensor(TAO)$313.332.77%
  • crypto-com-chainCronos(CRO)$0.07129010.21%
  • BitwayBitway(BTW)$1.309.53%
  • shiba-inuShiba Inu(SHIB)$0.0000062.09%
  • tether-goldTether Gold(XAUT)$4,157.750.04%
  • Global DollarGlobal Dollar(USDG)$1.000.02%
  • paypal-usdPayPal USD(PYUSD)$1.000.00%
  • aaveAave(AAVE)$167.3512.90%
  • okbOKB(OKB)$121.062.98%
  • EthenaEthena(ENA)$0.250913-4.17%
  • OndoOndo(ONDO)$0.52-0.03%
  • Ripple USDRipple USD(RLUSD)$1.00-0.01%
  • MemeCoreMemeCore(M)$1.07-9.02%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • Pump.funPump.fun(PUMP)$0.0050474.10%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.150.00%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Google Research Open-Sources RRSI: AI Agents That Improve Their Own Harness Without Overfitting

September 29, 2026
in AI & Technology
Reading Time: 6 mins read
A A
Google Research Open-Sources RRSI: AI Agents That Improve Their Own Harness Without Overfitting
ShareShareShareShareShare

Google Cloud AI Research, with UNC-Chapel Hill, Stanford and Washington University in St. Louis, has released RRSI (Regularized Recursive Self-Improvement). It lets an LLM agent rewrite its own harness: prompts, tools, memory, control flow and sub-agents. Model weights never change. RRSI constrains the improvement loop itself, so gains hold on benchmarks the agent never optimized against.

Deployable? Yes, as a research framework. The code is Apache 2.0, needs Python 3.10+, and accepts any LiteLLM model string. Defaults assume Claude Opus 4.8 on Vertex AI.

Why Self-Improving Harnesses Overfit

Harness evolution loops propose edits, score them on a fixed evolve set and keep the winner. The same tasks are reused every round, so the loop can memorize them. The RRSI research names 3 failure modes: benchmark-specific fitting, noise chasing and complexity accumulation. Each one widens the gap between evolve-set scores and real transfer.

How RRSI Works

RRSI keeps every harness component editable. It regularizes how the search moves instead.

Proposal side

  • Annealed edit budget: a cosine schedule lets early rounds bundle several edits. Late rounds allow a single attributable change.
  • Evidence-aware credit: each candidate is logged with its component, hypothesis, diff, score change and cost change. The proposer reads this ledger, so falsified ideas are not retried.
  • Structured exploration: when progress stalls inside the noise band, budget shifts to components the run never touched.

Selection side

  • Leakage critic: rejects task names, entities, answers or benchmark-specific logic before any scoring.
  • Noise-adjusted floor: gains must clear the variance measured on the unchanged base harness.
  • Cost rule: extra inference tokens must be paid for by measured gain.
  • Pruning: components that stop producing gains become deletion targets.

The research team frame these as analogies to classic regularizers. The edit budget maps to L0, pruning to Lasso (L1) and the cost rule to Ridge (L2).

Results Across 8 Benchmarks

All 6 held-out splits improved. With Gemini 3.5 Flash as the policy, Terminal-Bench 2.1 rose from 64.6 to 78.7. SWE-bench Verified rose from 76.8 to 79.0.

The harness is also lighter. On the agentic workspace instance, RRSI uses 2.42M policy tokens per trial. Unregularized evolution uses 3.80M. The abstract reports this as 30% fewer; the project page says 36%.

RRSI vs Closest Competitors

Scores come from Table 1 of the RRSI research paper. All methods share the same starting harness, policy, evolve split and candidate budget.

Feature RRSI Meta-Harness AHE TTHE HarnessX
Core idea Regularized proposal and selection Agentic proposer over code, scores and traces of all prior candidates Observability-driven loop; edits paired with verified predictions Evolves harness during test time, no gold labels Modular typed primitives, trace-driven adaptation
Model weights Frozen Frozen Frozen Frozen Frozen
Cost rule and pruning Yes No* No* No* No*
Harvey LAB evolve score 90.5 93.0 90.7 91.1 91.8
OOD average (H0 = 39.7) 43.6 40.6 39.2 38.0 39.7

*Per the RRSI research team. OOD average is the mean of JobBench, GDPval and APEX-Agents, computed from Table 1.

Meta-Harness leads on the Harvey LAB evolve split. RRSI has the smallest evolve gain but the only OOD average more than 1 point above H0.

Interactive Explainer

How RRSI Regularizes Agent Self-Improvement

The model stays frozen. The harness (prompts, tools, memory, control flow) evolves, but every edit must pass the regularizers.




Pick a candidate edit, then press Run. Watch where RRSI stops it.






ProposerReads full edit ledger, shrinking edit budget

Leakage criticScreens diff before scoring

EvaluateRun on evolve set

GateNoise floor + cost rule

Harness Ht+1Accepted, then pruning check

Candidates are illustrative. The rules they hit are the ones described in the RRSI paper.

// edit ledger: component | hypothesis | Δscore | Δcost | verdict



YOU MAY ALSO LIKE

Nothing’s Flagship $399 Headphone 1 Pro Actually Have Some Professional Features

OpenAI Reportedly Cancels GPT-6.1 Astra’s Release Over Deceptive Behavior

Getting Started

git clone https://github.com/google-research/rrsi.git && cd rrsi
pip install -e ".[dev]"
python3 rrsi.py --domain coding baseline
python3 rrsi.py --domain coding run

Each round drafts 2 candidates in separate git worktrees, screens them, evaluates both and fast-forwards the branch to the winner. The coding instance also needs Docker and harbor. New domains plug in through a single adapter module.

Key Takeaways

  • RRSI evolves prompts, tools, memory and workflows while model weights stay frozen.
  • A leakage critic, noise floor, cost rule and pruning decide which edits stick.
  • Terminal-Bench 2.1 rose from 74.2% to 80.2% with Claude Opus 4.8.
  • SWE-bench Verified, never used for selection, rose from 82.0% to 83.8%.
  • Apache 2.0 code on GitHub; research-grade, not an official Google product.

Check out the Paper, Codes and Technical details. All credit goes to the researcher of this project. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well.

Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.? Connect with us


Asif Razzaq is the CEO of Marktechpost AI Media Inc.. As a visionary entrepreneur and engineer, Asif is committed to harnessing the potential of Artificial Intelligence for social good. His most recent endeavor is the launch of an Artificial Intelligence Media Platform, Marktechpost, which stands out for its in-depth coverage of machine learning and deep learning news that is both technically sound and easily understandable by a wide audience. The platform boasts of over 2 million monthly views, illustrating its popularity among audiences.

Credit: Source link

ShareTweetSendSharePin

Related Posts

Nothing’s Flagship 9 Headphone 1 Pro Actually Have Some Professional Features
AI & Technology

Nothing’s Flagship $399 Headphone 1 Pro Actually Have Some Professional Features

September 29, 2026
OpenAI Reportedly Cancels GPT-6.1 Astra’s Release Over Deceptive Behavior
AI & Technology

OpenAI Reportedly Cancels GPT-6.1 Astra’s Release Over Deceptive Behavior

September 29, 2026
H Company Releases Holo4: Open-Weight Computer-Use Models That Click, Code and Call Tools Across Desktop, Web, Android and APIs
AI & Technology

H Company Releases Holo4: Open-Weight Computer-Use Models That Click, Code and Call Tools Across Desktop, Web, Android and APIs

September 29, 2026
Alibaba Qwen Releases Qwen-Audio-3.1-Realtime: A Full-Duplex Voice Model Trained to Think, Act, and Decide When to Speak
AI & Technology

Alibaba Qwen Releases Qwen-Audio-3.1-Realtime: A Full-Duplex Voice Model Trained to Think, Act, and Decide When to Speak

September 29, 2026
Next Post
Anthropic’s IPO prospectus shows sweeping AI vision and surging costs

Anthropic's IPO prospectus shows sweeping AI vision and surging costs

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Your Old GPU Could Be Worth More Than You Think

Your Old GPU Could Be Worth More Than You Think

September 26, 2026
Classic Bay Area tourist attraction closing after nearly 60 years

Classic Bay Area tourist attraction closing after nearly 60 years

September 23, 2026
Multiple Pullbacks Ahead — Kevin Mahn On What’s Worth Buying

Multiple Pullbacks Ahead — Kevin Mahn On What’s Worth Buying

September 23, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!