• bitcoinBitcoin(BTC)$85,466.004.52%
  • ethereumEthereum(ETH)$2,732.012.39%
  • tetherTether(USDT)$1.000.01%
  • binancecoinBNB(BNB)$787.501.80%
  • rippleXRP(XRP)$1.525.95%
  • usd-coinUSDC(USDC)$1.000.01%
  • solanaSolana(SOL)$116.754.05%
  • tronTRON(TRX)$0.3489451.67%
  • zcashZcash(ZEC)$1,483.58-1.68%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.011.28%
  • HyperliquidHyperliquid(HYPE)$94.20-0.10%
  • dogecoinDogecoin(DOGE)$0.10006912.44%
  • moneroMonero(XMR)$574.15-5.55%
  • whitebitWhiteBIT Coin(WBT)$85.932.86%
  • RainRain(RAIN)$0.013668-2.48%
  • chainlinkChainlink(LINK)$12.912.72%
  • USDSUSDS(USDS)$1.00-0.01%
  • cardanoCardano(ADA)$0.2448154.81%
  • leo-tokenLEO Token(LEO)$8.950.41%
  • stellarStellar(XLM)$0.2123366.44%
  • nearNEAR Protocol(NEAR)$4.433.76%
  • uniswapUniswap(UNI)$8.933.06%
  • bitcoin-cashBitcoin Cash(BCH)$265.294.28%
  • Ethena USDeEthena USDe(USDE)$1.00-0.03%
  • avalanche-2Avalanche(AVAX)$10.75-3.92%
  • litecoinLitecoin(LTC)$60.954.05%
  • CantonCanton(CC)$0.1183064.93%
  • daiDai(DAI)$1.000.00%
  • USD1USD1(USD1)$1.00-0.04%
  • suiSui(SUI)$1.025.76%
  • hedera-hashgraphHedera(HBAR)$0.0928106.01%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.442.39%
  • BittensorBittensor(TAO)$322.5418.55%
  • shiba-inuShiba Inu(SHIB)$0.0000068.20%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • crypto-com-chainCronos(CRO)$0.0657185.94%
  • MemeCoreMemeCore(M)$1.37-8.58%
  • paypal-usdPayPal USD(PYUSD)$1.000.01%
  • tether-goldTether Gold(XAUT)$4,320.58-0.70%
  • okbOKB(OKB)$121.271.46%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.140.01%
  • aaveAave(AAVE)$142.651.93%
  • EthenaEthena(ENA)$0.212706-1.65%
  • BitwayBitway(BTW)$0.7910.45%
  • pepePepe(PEPE)$0.00000526.19%
  • mantleMantle(MNT)$0.645.23%
  • Pump.funPump.fun(PUMP)$0.0045283.08%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Evaluating the Robustness and Fairness of Instruction-Tuned LLMs in Clinical Tasks: Implications for Performance Variability and Demographic Fairness

July 21, 2024
in AI & Technology
Reading Time: 4 mins read
A A
Evaluating the Robustness and Fairness of Instruction-Tuned LLMs in Clinical Tasks: Implications for Performance Variability and Demographic Fairness
ShareShareShareShareShare

Instruction-tuned LLMs can handle various tasks using natural language instructions, but their performance is sensitive to how instructions are phrased. This issue is critical in healthcare, where clinicians, who may need to be more skilled, prompt engineers, need reliable outputs. The robustness of LLMs to variations in clinical task instructions is thus questioned. Despite advancements in zero-shot task execution by models like GPT-3.5+, FLAN, Alpaca, and Mistral, their sensitivity to instruction phrasing poses challenges, especially in specialized domains like medicine, where inconsistent model performance can have significant consequences for patient care.

Researchers from Northeastern University and Codametrix collected prompts from medical doctors across various tasks to evaluate the sensitivity of seven general and specialized LLMs to natural instruction phrasings. They found substantial performance variability across all models, with domain-specific models trained on clinical data being particularly brittle. Differences in phrasing also affected fairness, with performance discrepancies observed between demographic groups in tasks like mortality prediction. The study highlights the robustness challenge in clinical LLMs and its implications for fairness, emphasizing the need for further research in this area. The researchers have released their code and prompts to support ongoing investigations.

YOU MAY ALSO LIKE

SpaceXAI Releases Grok 4.7: A Larger Base Model at the Same $2/$6 Price as Grok 4.6

Why It’s Important To Unplug Your PC During A Power Outage

Instruction-following LLMs have been enhanced to solve various tasks with minimal examples or instructions, thanks to techniques like Reinforcement Learning from Human Feedback and fine-tuning with labeled data. Large datasets for instruction tuning, such as Flan 2021 and Super-NaturalInstructions, have been created. However, LLMs are sensitive to prompt construction, affecting performance in few-shot and zero-shot settings. General LLMs can handle clinical tasks, but smaller, fine-tuned models often perform better. Privacy issues limit high-quality clinical datasets, leading researchers to use synthetic data, though general models usually outperform specialized ones. 

The study examines the robustness of LLMs to natural variations in instructional phrasings for clinical tasks. It involves ten clinical classification tasks and six information extraction tasks using data from MIMIC-III, i2b2, and n2c2 challenges. A diverse group of medical professionals wrote prompts for each task. Seven LLMs, including general-domain and domain-specific models, were evaluated for performance, variance, and fairness across these prompts. Models were assessed using zero-shot inference with specific sequence lengths, processing notes in chunks. AUROC scores for classification tasks and F1 scores for extraction tasks were reported to measure effectiveness.

Results for Mortality Prediction and Drug Extraction reveal significant variability in performance due to different yet semantically equivalent instructions. In Mortality Prediction, LLAMA 2 (13B) outperformed other models, while MISTRAL showed superior performance in other classification tasks. In Drug Extraction, LLAMA 2 (7B) performed best on average, though clinical models exhibited mixed results. Analysis of demographic subgroup performance indicated disparities, with non-White and female patients often receiving lower predictive accuracy. This variability in prompts affects fairness, highlighting that minor changes in phrasing can disproportionately impact certain demographic groups. General domain models generally outperformed clinical models across tasks.

In conclusion, the study evaluates instruction-tuned open-source LLMs for clinical classification and information extraction tasks on EHR clinical notes, focusing on robustness to variations in prompts from medical professionals. Twelve practitioners from diverse backgrounds wrote prompts for 16 clinical functions. Key findings include: LLM performance varies significantly across prompts from different experts; domain-specific models generally underperform compared to general models; and prompt variations impact fairness, leading to varying levels of fairness in outcomes. Practitioners should be cautious with instruction-tuned LLMs in critical clinical tasks, as minor phrasing differences can significantly affect outputs. This highlights the need for improved LLM robustness.


Check out the Paper. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter and join our Telegram Channel and LinkedIn Group. If you like our work, you will love our newsletter..

Don’t Forget to join our 46k+ ML SubReddit

Find Upcoming AI Webinars here


Sana Hassan, a consulting intern at Marktechpost and dual-degree student at IIT Madras, is passionate about applying technology and AI to address real-world challenges. With a keen interest in solving practical problems, he brings a fresh perspective to the intersection of AI and real-life solutions.

🐝 Join the Fastest Growing AI Research Newsletter Read by Researchers from Google + NVIDIA + Meta + Stanford + MIT + Microsoft and many others…


Credit: Source link

ShareTweetSendSharePin

Related Posts

SpaceXAI Releases Grok 4.7: A Larger Base Model at the Same / Price as Grok 4.6
AI & Technology

SpaceXAI Releases Grok 4.7: A Larger Base Model at the Same $2/$6 Price as Grok 4.6

September 22, 2026
Why It’s Important To Unplug Your PC During A Power Outage
AI & Technology

Why It’s Important To Unplug Your PC During A Power Outage

September 22, 2026
Why Is Your Laptop Fan So Loud?
AI & Technology

Why Is Your Laptop Fan So Loud?

September 22, 2026
These Drones Could Cover Up To 98 Percent Of The World’s Oceans By 2028
AI & Technology

These Drones Could Cover Up To 98 Percent Of The World’s Oceans By 2028

September 21, 2026
Next Post
‘The world is watching’ upcoming U.S. election ‘with a lot of fear and trepidation’: Full Panel

'The world is watching' upcoming U.S. election 'with a lot of fear and trepidation': Full Panel

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
At least 7 dead after ferry carrying nearly 270 capsizes off northern Cyprus

At least 7 dead after ferry carrying nearly 270 capsizes off northern Cyprus

September 21, 2026
Deadly shooting at gay bar investigated as hate crime

Deadly shooting at gay bar investigated as hate crime

September 19, 2026
Katie Stein, CEO of ASAPP – Interview Series – Unite.AI

Katie Stein, CEO of ASAPP – Interview Series – Unite.AI

September 15, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!