• bitcoinBitcoin(BTC)$83,246.00-1.36%
  • ethereumEthereum(ETH)$2,645.61-1.89%
  • tetherTether(USDT)$1.000.00%
  • binancecoinBNB(BNB)$768.34-0.43%
  • rippleXRP(XRP)$1.49-1.88%
  • usd-coinUSDC(USDC)$1.000.00%
  • solanaSolana(SOL)$119.65-0.79%
  • tronTRON(TRX)$0.3335320.22%
  • zcashZcash(ZEC)$1,559.42-5.02%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.06-0.38%
  • HyperliquidHyperliquid(HYPE)$89.61-3.90%
  • dogecoinDogecoin(DOGE)$0.093891-2.35%
  • chainlinkChainlink(LINK)$13.94-1.06%
  • moneroMonero(XMR)$540.31-3.44%
  • whitebitWhiteBIT Coin(WBT)$83.08-1.21%
  • USDSUSDS(USDS)$1.000.00%
  • cardanoCardano(ADA)$0.249869-0.77%
  • RainRain(RAIN)$0.012521-1.99%
  • leo-tokenLEO Token(LEO)$9.040.07%
  • stellarStellar(XLM)$0.210180-1.96%
  • nearNEAR Protocol(NEAR)$5.242.98%
  • bitcoin-cashBitcoin Cash(BCH)$313.48-5.60%
  • uniswapUniswap(UNI)$9.30-4.47%
  • CantonCanton(CC)$0.1405352.89%
  • litecoinLitecoin(LTC)$70.59-1.01%
  • suiSui(SUI)$1.245.72%
  • Ethena USDeEthena USDe(USDE)$1.00-0.01%
  • avalanche-2Avalanche(AVAX)$10.67-0.58%
  • daiDai(DAI)$1.00-0.02%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.623.60%
  • USD1USD1(USD1)$1.000.00%
  • hedera-hashgraphHedera(HBAR)$0.0944471.24%
  • quant-networkQuant(QNT)$250.7945.44%
  • BittensorBittensor(TAO)$307.32-4.28%
  • BitwayBitway(BTW)$1.2620.78%
  • shiba-inuShiba Inu(SHIB)$0.000006-2.83%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • crypto-com-chainCronos(CRO)$0.064289-3.11%
  • tether-goldTether Gold(XAUT)$4,198.77-1.88%
  • EthenaEthena(ENA)$0.2720661.99%
  • OndoOndo(ONDO)$0.565.51%
  • paypal-usdPayPal USD(PYUSD)$1.000.00%
  • MemeCoreMemeCore(M)$1.18-4.24%
  • Ripple USDRipple USD(RLUSD)$1.000.01%
  • okbOKB(OKB)$118.32-2.17%
  • Circle USYCCircle USYC(USYC)$1.140.00%
  • Pump.funPump.fun(PUMP)$0.00510016.14%
  • aaveAave(AAVE)$150.27-3.16%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.150.08%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Google Proposes TUMIX: Multi-Agent Test-Time Scaling With Tool-Use Mixture

October 4, 2025
in AI & Technology
Reading Time: 9 mins read
A A
Google Proposes TUMIX: Multi-Agent Test-Time Scaling With Tool-Use Mixture
ShareShareShareShareShare

What if, instead of re-sampling one agent, you could push Gemini-2.5 Pro to 34.1% on HLE by mixing 12–15 tool-using agents that share notes and stop early? Google Cloud AI Research, with collaborators from MIT, Harvard, and Google DeepMind, introduced TUMIX (Tool-Use Mixture)—a test-time framework that ensembles heterogeneous agent styles (text-only, code, search, guided variants) and lets them share intermediate answers over a few refinement rounds, then stop early via an LLM-based judge. The result: higher accuracy at lower cost on hard reasoning benchmarks such as HLE, GPQA-Diamond, and AIME (2024/2025).

https://arxiv.org/pdf/2510.01279

So, What exactly is different new?

  • Mixture over modality, not just more samples: TUMIX runs ~15 agent styles spanning Chain-of-Thought (CoT), code execution, web search, dual-tool agents, and guided variants. Each round, every agent sees (a) the original question and (b) other agents’ previous answers, then proposes a refined answer. This message-passing raises average accuracy early while diversity gradually collapses—so stopping matters.
  • Adaptive early-termination: An LLM-as-Judge halts refinement once answers exhibit strong consensus (with a minimum round threshold). This preserves accuracy at ~49% of the inference cost vs. fixed-round refinement; token cost drops to ~46% because late rounds are token-heavier.
  • Auto-designed agents: Beyond human-crafted agents, TUMIX prompts the base LLM to generate new agent types; mixing these with the manual set yields an additional ~+1.2% average lift without extra cost. The empirical “sweet spot” is ~12–15 agent styles.
https://arxiv.org/pdf/2510.01279

How does it work?

TUMIX runs a group of heterogeneous agents—text-only Chain-of-Thought, code-executing, web-searching, and guided variants—in parallel, then iterates a small number of refinement rounds where each agent conditions on the original question plus the other agents’ prior rationales and answers (structured note-sharing). After each round, an LLM-based judge evaluates consensus/consistency to decide early termination; if confidence is insufficient, another round is triggered, otherwise the system finalizes via simple aggregation (e.g., majority vote or selector). This mixture-of-tool-use design trades brute-force re-sampling for diverse reasoning paths, improving coverage of correct candidates while controlling token/tool budgets; empirically, benefits saturate around 12–15 agent styles, and stopping early preserves diversity and lowers cost without sacrificing accuracy

YOU MAY ALSO LIKE

20 Agentic Use Cases of TypeSafe AI’s Jev

Which Is Better To Use?

Lets discuss the Results

Under comparable inference budgets to strong tool-augmented baselines (Self-MoA, Symbolic-MoE, DEI, SciMaster, GSA), TUMIX yields the best average accuracy; a scaled variant (TUMIX+) pushes further with more compute:

  • HLE (Humanity’s Last Exam): Pro: 21.6% → 34.1% (TUMIX+); Flash: 9.7% → 23.1%.
    (HLE is a 2,500-question, difficult, multi-domain benchmark finalized in 2025.)
  • GPQA-Diamond: Pro: up to 88.3%; Flash: up to 82.1%. (GPQA-Diamond is the hardest 198-question subset authored by domain experts.)
  • AIME 2024/25: Pro: 96.7%; Flash: 86.7% with TUMIX(+) at test time.

Across tasks, TUMIX averages +3.55% over the best prior tool-augmented test-time scaling baseline at similar cost, and +7.8% / +17.4% over no-scaling for Pro/Flash, respectively.

https://arxiv.org/pdf/2510.01279

TUMIX is a great approach from Google because it frames test-time scaling as a search problem over heterogeneous tool policies rather than brute-force sampling. The parallel committee (text, code, search) improves candidate coverage, while the LLM-judge enables early-stop that preserves diversity and reduces token/tool spend—useful under latency budgets. The HLE gains (34.1% with Gemini-2.5 Pro) align with the benchmark’s finalized 2,500-question design, and the ~12–15 agent styles “sweet spot” indicates selection—not generation—is the limiting factor.


Check out the Paper. Feel free to check out our GitHub Page for Tutorials, Codes and Notebooks. Also, feel free to follow us on Twitter and don’t forget to join our 100k+ ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well.


Asif Razzaq is the CEO of Marktechpost Media Inc.. As a visionary entrepreneur and engineer, Asif is committed to harnessing the potential of Artificial Intelligence for social good. His most recent endeavor is the launch of an Artificial Intelligence Media Platform, Marktechpost, which stands out for its in-depth coverage of machine learning and deep learning news that is both technically sound and easily understandable by a wide audience. The platform boasts of over 2 million monthly views, illustrating its popularity among audiences.

🙌 Follow MARKTECHPOST: Add us as a preferred source on Google.

Credit: Source link

ShareTweetSendSharePin

Related Posts

20 Agentic Use Cases of TypeSafe AI’s Jev
AI & Technology

20 Agentic Use Cases of TypeSafe AI’s Jev

September 28, 2026
Which Is Better To Use?
AI & Technology

Which Is Better To Use?

September 28, 2026
Are 3D Printers Worth Buying In 2026?
AI & Technology

Are 3D Printers Worth Buying In 2026?

September 28, 2026
Bill Gates Says It’s ‘Completely Irresponsible’ For AI To Not Have Safeguards
AI & Technology

Bill Gates Says It’s ‘Completely Irresponsible’ For AI To Not Have Safeguards

September 27, 2026
Next Post
Assata Shakur, Black Liberation Army figure and activist, dies at 78

Assata Shakur, Black Liberation Army figure and activist, dies at 78

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
CNN staffers panic over Paramount merger as layoffs loom: ‘Bloodbath coming’

CNN staffers panic over Paramount merger as layoffs loom: ‘Bloodbath coming’

September 22, 2026
AI Trade Spreads, But Chips Were First

AI Trade Spreads, But Chips Were First

September 21, 2026
JEPQ: Hedging AI Infrastructure Growth With Monthly Cash Flow (NASDAQ:JEPQ)

JEPQ: Hedging AI Infrastructure Growth With Monthly Cash Flow (NASDAQ:JEPQ)

September 23, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!