• bitcoinBitcoin(BTC)$83,351.00-1.87%
  • ethereumEthereum(ETH)$2,681.06-1.17%
  • tetherTether(USDT)$1.00-0.01%
  • binancecoinBNB(BNB)$768.15-1.67%
  • rippleXRP(XRP)$1.52-1.51%
  • usd-coinUSDC(USDC)$1.00-0.01%
  • solanaSolana(SOL)$119.52-3.58%
  • tronTRON(TRX)$0.3355410.36%
  • zcashZcash(ZEC)$1,584.37-4.56%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.060.00%
  • HyperliquidHyperliquid(HYPE)$89.99-3.29%
  • dogecoinDogecoin(DOGE)$0.094294-3.96%
  • chainlinkChainlink(LINK)$14.833.53%
  • moneroMonero(XMR)$531.99-4.36%
  • whitebitWhiteBIT Coin(WBT)$83.37-1.67%
  • USDSUSDS(USDS)$1.00-0.04%
  • cardanoCardano(ADA)$0.252821-1.63%
  • RainRain(RAIN)$0.012544-1.24%
  • leo-tokenLEO Token(LEO)$9.01-0.12%
  • stellarStellar(XLM)$0.2245823.25%
  • nearNEAR Protocol(NEAR)$5.17-0.60%
  • bitcoin-cashBitcoin Cash(BCH)$313.04-7.53%
  • uniswapUniswap(UNI)$9.05-8.51%
  • litecoinLitecoin(LTC)$72.250.67%
  • hedera-hashgraphHedera(HBAR)$0.12281728.96%
  • CantonCanton(CC)$0.131870-4.40%
  • avalanche-2Avalanche(AVAX)$10.61-3.62%
  • Ethena USDeEthena USDe(USDE)$1.000.00%
  • suiSui(SUI)$1.20-5.45%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.651.98%
  • daiDai(DAI)$1.00-0.01%
  • USD1USD1(USD1)$1.000.00%
  • BittensorBittensor(TAO)$307.31-8.00%
  • BitwayBitway(BTW)$1.2812.25%
  • quant-networkQuant(QNT)$236.8244.14%
  • crypto-com-chainCronos(CRO)$0.0684330.78%
  • shiba-inuShiba Inu(SHIB)$0.000006-3.65%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • tether-goldTether Gold(XAUT)$4,163.12-2.74%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.01%
  • EthenaEthena(ENA)$0.268122-1.97%
  • MemeCoreMemeCore(M)$1.17-3.03%
  • OndoOndo(ONDO)$0.54-1.16%
  • Ripple USDRipple USD(RLUSD)$1.000.02%
  • okbOKB(OKB)$118.24-3.23%
  • Circle USYCCircle USYC(USYC)$1.140.00%
  • Pump.funPump.fun(PUMP)$0.00514413.73%
  • aaveAave(AAVE)$149.86-3.91%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.150.23%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Stanford Researchers Introduced MedAgentBench: A Real-World Benchmark for Healthcare AI Agents

September 16, 2025
in AI & Technology
Reading Time: 7 mins read
A A
Stanford Researchers Introduced MedAgentBench: A Real-World Benchmark for Healthcare AI Agents
ShareShareShareShareShare

A team of Stanford University researchers have released MedAgentBench, a new benchmark suite designed to evaluate large language model (LLM) agents in healthcare contexts. Unlike prior question-answering datasets, MedAgentBench provides a virtual electronic health record (EHR) environment where AI systems must interact, plan, and execute multi-step clinical tasks. This marks a significant shift from testing static reasoning to assessing agentic capabilities in live, tool-based medical workflows.

https://ai.nejm.org/doi/full/10.1056/AIdbp2500144

Why Do We Need Agentic Benchmarks in Healthcare?

Recent LLMs have moved beyond static chat-based interactions toward agentic behavior—interpreting high-level instructions, calling APIs, integrating patient data, and automating complex processes. In medicine, this evolution could help address staff shortages, documentation burden, and administrative inefficiencies.

YOU MAY ALSO LIKE

A Modular, Repairable GPS Watch Is A Good First Step

Fireworks AI Releases Ember-1: A Post-Trained Kimi K3 That Uses About 40% Fewer Tokens

While general-purpose agent benchmarks (e.g., AgentBench, AgentBoard, tau-bench) exist, healthcare lacked a standardized benchmark that captures the complexity of medical data, FHIR interoperability, and longitudinal patient records. MedAgentBench fills this gap by offering a reproducible, clinically relevant evaluation framework.

What Does MedAgentBench Contain?

How Are the Tasks Structured?

MedAgentBench consists of 300 tasks across 10 categories, written by licensed physicians. These tasks include patient information retrieval, lab result tracking, documentation, test ordering, referrals, and medication management. Tasks average 2–3 steps and mirror workflows encountered in inpatient and outpatient care.

What Patient Data Supports the Benchmark?

The benchmark leverages 100 realistic patient profiles extracted from Stanford’s STARR data repository, comprising over 700,000 records including labs, vitals, diagnoses, procedures, and medication orders. Data was de-identified and jittered for privacy while preserving clinical validity.

How Is the Environment Built?

The environment is FHIR-compliant, supporting both retrieval (GET) and modification (POST) of EHR data. AI systems can simulate realistic clinical interactions such as documenting vitals or placing medication orders. This design makes the benchmark directly translatable to live EHR systems.

How Are Models Evaluated?

  • Metric: Task success rate (SR), measured with strict pass@1 to reflect real-world safety requirements.
  • Models Tested: 12 leading LLMs including GPT-4o, Claude 3.5 Sonnet, Gemini 2.0, DeepSeek-V3, Qwen2.5, and Llama 3.3.
  • Agent Orchestrator: A baseline orchestration setup with nine FHIR functions, limited to eight interaction rounds per task.

Which Models Performed Best?

  • Claude 3.5 Sonnet v2: Best overall with 69.67% success, especially strong in retrieval tasks (85.33%).
  • GPT-4o: 64.0% success, showing balanced retrieval and action performance.
  • DeepSeek-V3: 62.67% success, leading among open-weight models.
  • Observation: Most models excelled at query tasks but struggled with action-based tasks requiring safe multi-step execution.
https://ai.nejm.org/doi/full/10.1056/AIdbp2500144

What Errors Did Models Make?

Two dominant failure patterns emerged:

  1. Instruction adherence failures — invalid API calls or incorrect JSON formatting.
  2. Output mismatch — providing full sentences when structured numerical values were required.

These errors highlight gaps in precision and reliability, both critical in clinical deployment.

Summary

MedAgentBench establishes the first large-scale benchmark for evaluating LLM agents in realistic EHR settings, pairing 300 clinician-authored tasks with a FHIR-compliant environment and 100 patient profiles. Results show strong potential but limited reliability—Claude 3.5 Sonnet v2 leads at 69.67%—highlighting the gap between query success and safe action execution. While constrained by single-institution data and EHR-focused scope, MedAgentBench provides an open, reproducible framework to drive the next generation of dependable healthcare AI agents


Check out the PAPER and Technical Blog. Feel free to check out our GitHub Page for Tutorials, Codes and Notebooks. Also, feel free to follow us on Twitter and don’t forget to join our 100k+ ML SubReddit and Subscribe to our Newsletter.


Michal Sutter is a data science professional with a Master of Science in Data Science from the University of Padova. With a solid foundation in statistical analysis, machine learning, and data engineering, Michal excels at transforming complex datasets into actionable insights.

Credit: Source link

ShareTweetSendSharePin

Related Posts

A Modular, Repairable GPS Watch Is A Good First Step
AI & Technology

A Modular, Repairable GPS Watch Is A Good First Step

September 28, 2026
Fireworks AI Releases Ember-1: A Post-Trained Kimi K3 That Uses About 40% Fewer Tokens
AI & Technology

Fireworks AI Releases Ember-1: A Post-Trained Kimi K3 That Uses About 40% Fewer Tokens

September 28, 2026
You Can Now Preorder The Tiny Boox Picco Ereader
AI & Technology

You Can Now Preorder The Tiny Boox Picco Ereader

September 28, 2026
20 Agentic Use Cases of TypeSafe AI’s Jev
AI & Technology

20 Agentic Use Cases of TypeSafe AI’s Jev

September 28, 2026
Next Post
Parents push for legislation amid new allegations that Meta suppressed child safety research

Parents push for legislation amid new allegations that Meta suppressed child safety research

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
FDA approves Eli Lilly’s Onswik once-weekly insulin injection for type 2 diabetes

FDA approves Eli Lilly’s Onswik once-weekly insulin injection for type 2 diabetes

September 24, 2026
Dust devil swirls through an Alabama softball game

Dust devil swirls through an Alabama softball game

September 24, 2026
Live updates: Trump says he faces a decision on Iran to make a deal or ‘annihilate’ the country – CNN

Live updates: Trump says he faces a decision on Iran to make a deal or ‘annihilate’ the country – CNN

September 22, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!