• bitcoinBitcoin(BTC)$77,554.001.38%
  • ethereumEthereum(ETH)$2,487.181.58%
  • tetherTether(USDT)$1.000.00%
  • binancecoinBNB(BNB)$755.574.17%
  • rippleXRP(XRP)$1.321.54%
  • usd-coinUSDC(USDC)$1.00-0.01%
  • solanaSolana(SOL)$105.715.62%
  • tronTRON(TRX)$0.3359480.20%
  • zcashZcash(ZEC)$1,490.499.38%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.03-0.15%
  • HyperliquidHyperliquid(HYPE)$87.7010.49%
  • dogecoinDogecoin(DOGE)$0.0842113.74%
  • moneroMonero(XMR)$532.617.21%
  • USDSUSDS(USDS)$1.000.02%
  • whitebitWhiteBIT Coin(WBT)$79.901.51%
  • RainRain(RAIN)$0.0130711.22%
  • chainlinkChainlink(LINK)$11.795.23%
  • leo-tokenLEO Token(LEO)$8.90-0.26%
  • cardanoCardano(ADA)$0.2135447.46%
  • stellarStellar(XLM)$0.1870502.14%
  • uniswapUniswap(UNI)$8.6526.81%
  • bitcoin-cashBitcoin Cash(BCH)$246.8511.32%
  • Ethena USDeEthena USDe(USDE)$1.00-0.02%
  • nearNEAR Protocol(NEAR)$3.5028.36%
  • daiDai(DAI)$1.000.02%
  • USD1USD1(USD1)$1.000.01%
  • CantonCanton(CC)$0.1090599.88%
  • litecoinLitecoin(LTC)$55.084.82%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.362.57%
  • avalanche-2Avalanche(AVAX)$7.894.66%
  • hedera-hashgraphHedera(HBAR)$0.0760362.70%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • suiSui(SUI)$0.787.71%
  • shiba-inuShiba Inu(SHIB)$0.0000055.72%
  • crypto-com-chainCronos(CRO)$0.0587040.36%
  • MemeCoreMemeCore(M)$1.2714.18%
  • paypal-usdPayPal USD(PYUSD)$1.000.01%
  • BittensorBittensor(TAO)$243.827.76%
  • tether-goldTether Gold(XAUT)$4,385.111.44%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • okbOKB(OKB)$114.072.13%
  • Ripple USDRipple USD(RLUSD)$1.000.01%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.15-0.06%
  • aaveAave(AAVE)$134.388.74%
  • AsterAster(ASTER)$0.752.35%
  • Pump.funPump.fun(PUMP)$0.00432411.91%
  • polkadotPolkadot(DOT)$1.1513.49%
  • mantleMantle(MNT)$0.594.78%
  • pax-goldPAX Gold(PAXG)$4,385.161.41%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

AgentClinic: Simulating Clinical Environments for Assessing Language Models in Healthcare

May 17, 2024
in AI & Technology
Reading Time: 5 mins read
A A
AgentClinic: Simulating Clinical Environments for Assessing Language Models in Healthcare
ShareShareShareShareShare

The primary goal of AI is to create interactive systems capable of solving diverse problems, including those in medical AI aimed at improving patient outcomes. Large language models (LLMs)  have demonstrated significant problem-solving abilities, surpassing human scores on exams like the USMLE. While LLMs can enhance healthcare accessibility, they still face limitations in real-world clinical settings due to the complexity of clinical tasks involving sequential decision-making, handling uncertainty, and compassionate patient care. Current evaluations mostly focus on static multiple-choice questions, not fully capturing the dynamic nature of clinical work.

The USMLE assesses medical students across foundational knowledge, clinical application, and independent practice skills. In contrast, the Objective Structured Clinical Examination (OSCE) evaluates practical clinical skills through simulated scenarios, offering direct observation and a comprehensive assessment. Language models in medicine are primarily evaluated using knowledge-based benchmarks like MedQA, which consists of challenging medical question-answering pairs. Recent efforts focus on refining language models’ applications in healthcare through red teaming and creating new benchmarks like EquityMedQA to address biases and improve evaluation methods. Also, advancements in clinical decision-making simulations, such as AMIE, show promise in enhancing diagnostic accuracy in medical AI.

Researchers from  Stanford University, Johns Hopkins University, and Hospital Israelita Albert Einstein present AgentClinic, an open-source benchmark for simulating clinical environments using language, patient, doctor, and measurement agents. It extends previous simulations by including medical exams (e.g., temperature, blood pressure) and ordering medical images (e.g., MRI, X-ray) through dialogue. Also, AgentClinic supports 24 biases found in clinical settings.

AgentClinic introduces four language agents: patient, doctor, measurement, and moderator. Each agent has specific roles and unique information for simulating clinical interactions. The patient agent provides symptom information without knowing the diagnosis, the measurement agent offers medical readings and test results, the doctor agent evaluates the patient and requests tests, and the moderator assesses the doctor’s diagnosis. AgentClinic also includes 24 biases relevant to clinical settings. The agents are built using curated medical questions from the USMLE and NEJM case challenges to create structured scenarios for evaluation using language models like GPT-4.

The accuracy of different language models (GPT-4, Mixtral-8x7B, GPT-3.5, and Llama 2 70B-chat) is evaluated on AgentClinic-MedQA, where each model acts as a doctor agent diagnosing patients through dialogue. GPT-4 achieved the highest accuracy at 52%, followed by GPT-3.5 at 38%, Mixtral-8x7B at 37%, and Llama 2 at 70B-chat at 9%. Comparison with MedQA accuracy showed weak predictability for AgentClinic-MedQA accuracy, similar to studies on medical residents’ performance relative to the USMLE.

To recapitulate,  this work researchers present AgentClinic, a benchmark for simulating clinical environments with 15 multimodal language agents and 107 unique language agents based on USMLE cases. These agents exhibit 23 biases, impacting diagnostic accuracy and patient-doctor interactions. GPT-4, the highest-performing model, shows reduced accuracy (1.7%-2%) with cognitive biases and larger reductions (1.5%) with implicit biases, affecting patient follow-up willingness and confidence. Cross-communication between patient and doctor models improves accuracy. Limited or excessive interaction time decreases accuracy, with a 27% reduction at N=10 interactions and a 4%-9% reduction at N>20 interactions. GPT-4V achieves around 27% accuracy in a multimodal clinical environment based on NEJM cases.


Check out the Paper and Project. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter. Join our Telegram Channel, Discord Channel, and LinkedIn Group.

If you like our work, you will love our newsletter..

Don’t Forget to join our 42k+ ML SubReddit


YOU MAY ALSO LIKE

Meta Launches Muse Mac App With File, Messages, and Calendar Access – Unite.AI

Waymo Announces Singapore Expansion Targeting 2028 Ride-Hailing Launch – Unite.AI

Aswin AK is a consulting intern at MarkTechPost. He is pursuing his Dual Degree at the Indian Institute of Technology, Kharagpur. He is passionate about data science and machine learning, bringing a strong academic background and hands-on experience in solving real-life cross-domain challenges.


🐝 Join the Fastest Growing AI Research Newsletter Read by Researchers from Google + NVIDIA + Meta + Stanford + MIT + Microsoft and many others…


Credit: Source link

ShareTweetSendSharePin

Related Posts

Meta Launches Muse Mac App With File, Messages, and Calendar Access – Unite.AI
AI & Technology

Meta Launches Muse Mac App With File, Messages, and Calendar Access – Unite.AI

September 18, 2026
Waymo Announces Singapore Expansion Targeting 2028 Ride-Hailing Launch – Unite.AI
AI & Technology

Waymo Announces Singapore Expansion Targeting 2028 Ride-Hailing Launch – Unite.AI

September 18, 2026
eGPUs Do Work, But They Come With Some Notable Limitations
AI & Technology

eGPUs Do Work, But They Come With Some Notable Limitations

September 17, 2026
Google’s Revamped CC Is An AI Agent For Families And Groups
AI & Technology

Google’s Revamped CC Is An AI Agent For Families And Groups

September 17, 2026
Next Post
Hidden cholesterol risk could affect millions of Americans

Hidden cholesterol risk could affect millions of Americans

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Special Report: America marks 25th anniversary of 9/11 attacks

Special Report: America marks 25th anniversary of 9/11 attacks

September 13, 2026
Reddington on the balance between empathy and the law

Reddington on the balance between empathy and the law

September 15, 2026
KSLV Vs. SLVP: A 26% Yield Hasn't Been Enough

KSLV Vs. SLVP: A 26% Yield Hasn't Been Enough

September 11, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!