• bitcoinBitcoin(BTC)$83,002.00-2.16%
  • ethereumEthereum(ETH)$2,662.70-1.65%
  • tetherTether(USDT)$1.00-0.01%
  • binancecoinBNB(BNB)$762.87-2.04%
  • rippleXRP(XRP)$1.49-3.03%
  • usd-coinUSDC(USDC)$1.000.00%
  • solanaSolana(SOL)$118.58-4.29%
  • tronTRON(TRX)$0.334122-0.04%
  • zcashZcash(ZEC)$1,564.32-5.85%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.060.00%
  • HyperliquidHyperliquid(HYPE)$89.85-3.44%
  • dogecoinDogecoin(DOGE)$0.093156-4.55%
  • chainlinkChainlink(LINK)$13.96-2.30%
  • moneroMonero(XMR)$529.95-4.54%
  • whitebitWhiteBIT Coin(WBT)$82.98-1.93%
  • USDSUSDS(USDS)$1.00-0.04%
  • cardanoCardano(ADA)$0.245537-3.95%
  • RainRain(RAIN)$0.012564-0.86%
  • leo-tokenLEO Token(LEO)$9.06-0.08%
  • stellarStellar(XLM)$0.211981-2.51%
  • nearNEAR Protocol(NEAR)$5.09-2.07%
  • bitcoin-cashBitcoin Cash(BCH)$310.19-8.10%
  • litecoinLitecoin(LTC)$71.680.34%
  • uniswapUniswap(UNI)$8.95-10.58%
  • CantonCanton(CC)$0.132374-4.10%
  • hedera-hashgraphHedera(HBAR)$0.11440220.86%
  • avalanche-2Avalanche(AVAX)$10.55-3.75%
  • Ethena USDeEthena USDe(USDE)$1.000.00%
  • suiSui(SUI)$1.18-5.04%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.672.84%
  • daiDai(DAI)$1.000.04%
  • USD1USD1(USD1)$1.00-0.01%
  • BittensorBittensor(TAO)$305.29-7.09%
  • quant-networkQuant(QNT)$237.8043.16%
  • shiba-inuShiba Inu(SHIB)$0.000006-4.80%
  • BitwayBitway(BTW)$1.228.16%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • crypto-com-chainCronos(CRO)$0.064727-6.08%
  • tether-goldTether Gold(XAUT)$4,157.85-2.85%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.02%
  • MemeCoreMemeCore(M)$1.17-3.39%
  • EthenaEthena(ENA)$0.262564-3.76%
  • OndoOndo(ONDO)$0.52-4.03%
  • Ripple USDRipple USD(RLUSD)$1.000.02%
  • okbOKB(OKB)$117.55-3.52%
  • Circle USYCCircle USYC(USYC)$1.140.00%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • aaveAave(AAVE)$148.06-4.76%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.150.11%
  • Pump.funPump.fun(PUMP)$0.0048799.69%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Meta AI Researchers Introduced SWEET-RL and CollaborativeAgentBench: A Step-Wise Reinforcement Learning Framework to Train Multi-Turn Language Agents for Realistic Human-AI Collaboration Tasks

March 23, 2025
in AI & Technology
Reading Time: 6 mins read
A A
Meta AI Researchers Introduced SWEET-RL and CollaborativeAgentBench: A Step-Wise Reinforcement Learning Framework to Train Multi-Turn Language Agents for Realistic Human-AI Collaboration Tasks
ShareShareShareShareShare

Large language models (LLMs) are rapidly transforming into autonomous agents capable of performing complex tasks that require reasoning, decision-making, and adaptability. These agents are deployed in web navigation, personal assistance, and software development. To act effectively in real-world settings, these agents must handle multi-turn interactions that span several steps or decision points. This introduces the need for training methods beyond simple response generation and instead focuses on optimizing the entire trajectory of interactions. Reinforcement learning (RL) has emerged as a compelling approach to train such agents by refining their decision-making based on long-term rewards.

Despite their potential, LLM-based agents struggle with multi-turn decision-making. A major challenge lies in assigning proper credit to actions taken at earlier stages of interaction, which influence later outcomes. Traditional training methods rely on next-token prediction or imitate high-probability actions, which do not account for long-term dependencies or cumulative goals. As a result, these methods fail to address the high variance and inefficiency of long-horizon tasks, particularly in collaborative scenarios where understanding human intent and reasoning across multiple steps is critical.

YOU MAY ALSO LIKE

Fireworks AI Releases Ember-1: A Post-Trained Kimi K3 That Uses About 40% Fewer Tokens

You Can Now Preorder The Tiny Boox Picco Ereader

Various reinforcement learning techniques have been adapted to fine-tune LLMs, especially from single-turn human feedback scenarios. Tools like PPO, RAFT, and DPO have been explored but exhibit significant limitations when applied to sequential interactions. These methods often fail at effective credit assignment across turns, making them less effective for multi-turn decision-making tasks. Benchmarks used to evaluate such tools lack the diversity and complexity required to assess performance in collaborative, real-world settings robustly. Value-based learning approaches are another alternative, but their need for custom heads and large amounts of task-specific fine-tuning data limit their generalization capabilities.

FAIR at Meta and UC Berkeley researchers proposed a new reinforcement learning method called SWEET-RL (Step-WisE Evaluation from Training-time Information). They also introduced a benchmark known as CollaborativeAgentBench or ColBench. This benchmark is central to the study, providing over 10,000 training tasks and over 1,000 test cases across two domains: backend programming and frontend design. ColBench simulates real collaboration between an AI agent and a human partner, where agents must ask questions, refine their understanding, and provide iterative solutions. For programming, agents are required to write functions in Python by asking for clarifications to refine missing specifications. In front-end tasks, agents must generate HTML code that matches a visual target through feedback-based corrections. Each task is designed to stretch the reasoning ability of the agent and mimic real-world constraints like limited interactions, capped at 10 turns per session.

SWEET-RL is built around an asymmetric actor-critic structure. The critic has access to additional information during training, such as the correct solution, which is not visible to the actor. This information allows the critic to evaluate each decision made by the agent with a much finer resolution. Instead of training a value function that estimates overall reward, SWEET-RL directly models an advantage function at each turn, using the Bradley-Terry optimization objective. The advantage function determines how much better or worse a particular action is compared to alternatives, helping the agent learn precise behaviors. For example, if an action aligns better with the human partner’s expectation, it receives a higher advantage score. This method simplifies credit assignment and aligns better with the pre-training architecture of LLMs, which rely on token-level prediction.

SWEET-RL achieved a 6% absolute improvement over other multi-turn reinforcement learning methods across both programming and design tasks. On backend programming tasks, it passed 48.0% of tests and achieved a success rate of 34.4%, compared to 28.2% for Multi-Turn DPO and 22.4% for zero-shot performance. On frontend design tasks, it reached a cosine similarity score of 76.9% and a win rate of 40.4%, improving from 38.6% with DPO and 33.8% with fine-tuning. Even when evaluated against top proprietary models like GPT-4o and O1-Mini, SWEET-RL closed the performance gap significantly, enabling the open-source Llama-3.1-8B model to match or exceed GPT-4o’s frontend win rate of 40.4%.

This research demonstrates that effective training of interactive agents hinges on precise, turn-by-turn feedback rather than generalized value estimations or broad supervision. SWEET-RL significantly improves credit assignment by leveraging training-time information and an architecture-aligned optimization approach. It enhances generalization, reduces training variance, and shows strong scalability, achieving better results with increased data. The algorithm also remains effective when applied to off-policy datasets, underlining its practicality in real-world scenarios with imperfect data. The research team created a meaningful evaluation framework by introducing ColBench as a benchmark tailored for realistic, multi-turn tasks. This combination with SWEET-RL provides a strong foundation for developing agents that can reason, adapt, and collaborate effectively over extended interactions.

Several key takeaways from this research include:

  1. SWEET-RL improved backend programming success rates from 28.2% (DPO) to 34.4% and frontend win rates from 38.6% to 40.4%.  
  2. It allowed Llama-3.1-8B to match the performance of GPT-4o, reducing dependency on proprietary models.  
  3. The critic uses training-time information (e.g., correct solutions) that is invisible to the actor, creating an asymmetric training setup.  
  4. Tasks in ColBench are capped at 10 rounds per session and include over 10,000 procedurally generated training examples.  
  5. ColBench measures outcomes using unit test pass rates (for code) and cosine similarity (for web design), providing reliable evaluation.  
  6. SWEET-RL directly learns a turn-wise advantage function, improving credit assignment without needing an intermediate value function.  
  7. The model scales effectively with more data and performs well even on off-policy datasets from weaker models.  
  8. Compared to traditional fine-tuning methods, SWEET-RL delivers higher performance with less overfitting and greater generalization.

Check out the Paper, GitHub Page and Dataset. All credit for this research goes to the researchers of this project. Also, feel free to follow us on Twitter and don’t forget to join our 85k+ ML SubReddit.


Nikhil is an intern consultant at Marktechpost. He is pursuing an integrated dual degree in Materials at the Indian Institute of Technology, Kharagpur. Nikhil is an AI/ML enthusiast who is always researching applications in fields like biomaterials and biomedical science. With a strong background in Material Science, he is exploring new advancements and creating opportunities to contribute.

Credit: Source link

ShareTweetSendSharePin

Related Posts

Fireworks AI Releases Ember-1: A Post-Trained Kimi K3 That Uses About 40% Fewer Tokens
AI & Technology

Fireworks AI Releases Ember-1: A Post-Trained Kimi K3 That Uses About 40% Fewer Tokens

September 28, 2026
You Can Now Preorder The Tiny Boox Picco Ereader
AI & Technology

You Can Now Preorder The Tiny Boox Picco Ereader

September 28, 2026
20 Agentic Use Cases of TypeSafe AI’s Jev
AI & Technology

20 Agentic Use Cases of TypeSafe AI’s Jev

September 28, 2026
Google Research Introduces an AI Video Co-Director: 4 Agentic Frameworks for Coherent, Minutes-Long Video Generation
AI & Technology

Google Research Introduces an AI Video Co-Director: 4 Agentic Frameworks for Coherent, Minutes-Long Video Generation

September 28, 2026
Next Post
Fin-R1: A Specialized Large Language Model for Financial Reasoning and Decision-Making

Fin-R1: A Specialized Large Language Model for Financial Reasoning and Decision-Making

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
AriZona Iced Tea billionaire allegedly promised masseuse an 0K home

AriZona Iced Tea billionaire allegedly promised masseuse an $820K home

September 22, 2026
I Borrowed 0,000 For Family And They Haven’t Paid Me Back

I Borrowed $300,000 For Family And They Haven’t Paid Me Back

September 22, 2026
Ex-NBA player Enes Kanter Freedom ejected from WNBA game after confronting Chicago’s Natasha Cloud

Ex-NBA player Enes Kanter Freedom ejected from WNBA game after confronting Chicago’s Natasha Cloud

September 25, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!