• bitcoinBitcoin(BTC)$85,739.006.33%
  • ethereumEthereum(ETH)$2,733.875.88%
  • tetherTether(USDT)$1.000.00%
  • binancecoinBNB(BNB)$795.975.60%
  • rippleXRP(XRP)$1.498.11%
  • usd-coinUSDC(USDC)$1.000.01%
  • solanaSolana(SOL)$117.918.89%
  • tronTRON(TRX)$0.344794-0.08%
  • zcashZcash(ZEC)$1,521.855.94%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.020.00%
  • HyperliquidHyperliquid(HYPE)$93.973.17%
  • dogecoinDogecoin(DOGE)$0.0935949.87%
  • moneroMonero(XMR)$576.076.47%
  • whitebitWhiteBIT Coin(WBT)$86.255.25%
  • RainRain(RAIN)$0.0140929.31%
  • chainlinkChainlink(LINK)$12.976.86%
  • USDSUSDS(USDS)$1.000.00%
  • cardanoCardano(ADA)$0.2421399.15%
  • leo-tokenLEO Token(LEO)$8.990.59%
  • stellarStellar(XLM)$0.2086568.85%
  • uniswapUniswap(UNI)$8.863.09%
  • bitcoin-cashBitcoin Cash(BCH)$267.778.88%
  • nearNEAR Protocol(NEAR)$4.1011.31%
  • avalanche-2Avalanche(AVAX)$11.192.76%
  • Ethena USDeEthena USDe(USDE)$1.000.01%
  • litecoinLitecoin(LTC)$62.489.34%
  • daiDai(DAI)$1.000.00%
  • CantonCanton(CC)$0.1147298.62%
  • USD1USD1(USD1)$1.000.01%
  • suiSui(SUI)$1.0425.67%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.434.27%
  • hedera-hashgraphHedera(HBAR)$0.0910313.72%
  • MemeCoreMemeCore(M)$1.50-3.02%
  • shiba-inuShiba Inu(SHIB)$0.0000067.43%
  • Global DollarGlobal Dollar(USDG)$1.000.01%
  • BittensorBittensor(TAO)$283.3312.39%
  • crypto-com-chainCronos(CRO)$0.0639529.40%
  • paypal-usdPayPal USD(PYUSD)$1.000.01%
  • tether-goldTether Gold(XAUT)$4,336.16-0.75%
  • okbOKB(OKB)$122.545.25%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • BitwayBitway(BTW)$0.9025.28%
  • Ripple USDRipple USD(RLUSD)$1.00-0.01%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.140.48%
  • EthenaEthena(ENA)$0.22296510.86%
  • aaveAave(AAVE)$145.157.93%
  • OndoOndo(ONDO)$0.4474189.15%
  • mantleMantle(MNT)$0.636.14%
  • AsterAster(ASTER)$0.763.46%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Stanford Researchers Released AgentFlow: In-the-Flow Reinforcement Learning RL for Modular, Tool-Using AI Agents

October 9, 2025
in AI & Technology
Reading Time: 6 mins read
A A
Stanford Researchers Released AgentFlow: In-the-Flow Reinforcement Learning RL for Modular, Tool-Using AI Agents
ShareShareShareShareShare

TL;DR: AgentFlow is a trainable agent framework with four modules—Planner, Executor, Verifier, Generator—coordinated by an explicit memory and toolset. The planner is optimized in the loop with a new on-policy method, Flow-GRPO, which broadcasts a trajectory-level outcome reward to every turn and applies token-level PPO-style updates with KL regularization and group-normalized advantages. On ten benchmarks, a 7B backbone tuned with Flow-GRPO reports +14.9% (search), +14.0% (agentic), +14.5% (math), and +4.1% (science) over strong baselines.

What is AgentFlow?

AgentFlow formalizes multi-turn, tool-integrated reasoning as an Markov Decision Process (MDP). At each turn, the Planner proposes a sub-goal and selects a tool plus context; the Executor calls the tool; the Verifier signals whether to continue; the Generator emits the final answer on termination. A structured, evolving memory records states, tool calls, and verification signals, constraining context growth and making trajectories auditable. Only the planner is trained; other modules can be fixed engines.

YOU MAY ALSO LIKE

A Laptop That Works Better With Your Android Phone

How AI Modernizes Lending Alongside Legacy Banking Systems Without a Teardown – Unite.AI

The public implementation showcases a modular toolkit (e.g., base_generator, python_coder, google_search, wikipedia_search, web_search) and ships quick-start scripts for inference, training, and benchmarking. The repository is MIT-licensed.

https://arxiv.org/pdf/2510.05592

Training method: Flow-GRPO

Flow-GRPO (Flow-based Group Refined Policy Optimization) converts long-horizon, sparse-reward optimization into tractable single-turn updates:

  • Final-outcome reward broadcast: a single, verifiable trajectory-level signal (LLM-as-judge correctness) is assigned to every turn, aligning local planning steps with global success.
  • Token-level clipped objective: importance-weighted ratios are computed per token, with PPO-style clipping and a KL penalty to a reference policy to prevent drift.
  • Group-normalized advantages: variance reduction across groups of on-policy rollouts stabilizes updates.
https://arxiv.org/pdf/2510.05592

Understanding the results and benchmarks

Benchmarks. The research team evaluates four task types: knowledge-intensive search (Bamboogle, 2Wiki, HotpotQA, Musique), agentic reasoning (GAIA textual split), math (AIME-24, AMC-23, Game of 24), and science (GPQA, MedQA). GAIA is a tooling-oriented benchmark for general assistants; the textual split excludes multimodal requirements.

Main numbers (7B backbone after Flow-GRPO). Average gains over strong baselines: +14.9% (search), +14.0% (agentic), +14.5% (math), +4.1% (science). The research team state their 7B system surpasses GPT-4o on the reported suite. The project page also reports training effects such as improved planning quality, reduced tool-calling errors (up to 28.4% on GAIA), and positive trends with larger turn budgets and model scale.

Ablations. Online Flow-GRPO improves performance by +17.2% vs. a frozen-planner baseline, while offline supervised fine-tuning of the planner degrades performance by −19.0% on their composite metric.

https://arxiv.org/pdf/2510.05592

Key Takeaways

  • Modular agent, planner-only training. AgentFlow structures an agent into Planner–Executor–Verifier–Generator with an explicit memory; only the Planner is trained in-loop.
  • Flow-GRPO converts long-horizon RL to single-turn updates. A trajectory-level outcome reward is broadcast to every turn; updates use token-level PPO-style clipping with KL regularization and group-normalized advantages.
  • The research team-reported gains on 10 benchmarks. With a 7B backbone, AgentFlow reports average improvements of +14.9% (search), +14.0% (agentic/GAIA textual), +14.5% (math), +4.1% (science) over strong baselines, and states surpassing GPT-4o on the same suite.
  • Tool-use reliability improves. The research team report reduced tool-calling errors (e.g., on GAIA) and better planning quality under larger turn budgets and model scale.

Editorial Comments

AgentFlow formalizes tool-using agents into four modules (planner, executor, verifier, generator) and trains only the planner in-loop via Flow-GRPO, which broadcasts a single trajectory-level reward to every turn with token-level PPO-style updates and KL control. Reported results on ten benchmarks show average gains of +14.9% (search), +14.0% (agentic/GAIA textual split), +14.5% (math), and +4.1% (science); the research team additionally state the 7B system surpasses GPT-4o on this suite. Implementation, tools, and quick-start scripts are MIT-licensed in the GitHub repo.


Check out the Technical Paper, GitHub Page and Project Page. Feel free to check out our GitHub Page for Tutorials, Codes and Notebooks. Also, feel free to follow us on Twitter and don’t forget to join our 100k+ ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well.

The post Stanford Researchers Released AgentFlow: In-the-Flow Reinforcement Learning RL for Modular, Tool-Using AI Agents appeared first on MarkTechPost.

Credit: Source link

ShareTweetSendSharePin

Related Posts

A Laptop That Works Better With Your Android Phone
AI & Technology

A Laptop That Works Better With Your Android Phone

September 21, 2026
How AI Modernizes Lending Alongside Legacy Banking Systems Without a Teardown – Unite.AI
AI & Technology

How AI Modernizes Lending Alongside Legacy Banking Systems Without a Teardown – Unite.AI

September 21, 2026
Best Voice Cloning APIs in 2026: Speaker Similarity, Consent Checks, and Price per 1M Characters
AI & Technology

Best Voice Cloning APIs in 2026: Speaker Similarity, Consent Checks, and Price per 1M Characters

September 21, 2026
StepFun Launches Step 5 Preview: A 600B-Total, 27B-Active MoE Model With 1M Context for Long-Horizon Agentic Work
AI & Technology

StepFun Launches Step 5 Preview: A 600B-Total, 27B-Active MoE Model With 1M Context for Long-Horizon Agentic Work

September 21, 2026
Next Post
Investigators believe remains of Travis Decker found

Investigators believe remains of Travis Decker found

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Protalix BioTherapeutics, Inc. (PLX) Presents at Morgan Stanley 24th Annual Global Healthcare Conference Transcript

Protalix BioTherapeutics, Inc. (PLX) Presents at Morgan Stanley 24th Annual Global Healthcare Conference Transcript

September 16, 2026
Oracle: OpenAI Just Blinked

Oracle: OpenAI Just Blinked

September 19, 2026
Spike on 10-year bond yields renews concerns over U.S. debt – The Washington Post

Spike on 10-year bond yields renews concerns over U.S. debt – The Washington Post

September 15, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!