• bitcoinBitcoin(BTC)$85,920.000.05%
  • ethereumEthereum(ETH)$2,729.78-0.37%
  • tetherTether(USDT)$1.000.00%
  • binancecoinBNB(BNB)$783.40-1.72%
  • rippleXRP(XRP)$1.553.33%
  • usd-coinUSDC(USDC)$1.000.00%
  • solanaSolana(SOL)$116.48-1.46%
  • tronTRON(TRX)$0.341596-0.92%
  • zcashZcash(ZEC)$1,517.77-0.16%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.02-0.07%
  • HyperliquidHyperliquid(HYPE)$94.10-0.11%
  • dogecoinDogecoin(DOGE)$0.0994015.84%
  • moneroMonero(XMR)$572.16-0.72%
  • whitebitWhiteBIT Coin(WBT)$86.370.01%
  • chainlinkChainlink(LINK)$12.87-1.05%
  • USDSUSDS(USDS)$1.000.00%
  • RainRain(RAIN)$0.013428-4.83%
  • cardanoCardano(ADA)$0.2487472.32%
  • leo-tokenLEO Token(LEO)$8.97-0.16%
  • stellarStellar(XLM)$0.2115521.13%
  • bitcoin-cashBitcoin Cash(BCH)$320.7419.50%
  • nearNEAR Protocol(NEAR)$4.345.26%
  • uniswapUniswap(UNI)$8.88-0.36%
  • Ethena USDeEthena USDe(USDE)$1.00-0.03%
  • avalanche-2Avalanche(AVAX)$11.01-1.64%
  • litecoinLitecoin(LTC)$61.27-2.02%
  • daiDai(DAI)$1.000.02%
  • CantonCanton(CC)$0.1156960.51%
  • USD1USD1(USD1)$1.00-0.03%
  • hedera-hashgraphHedera(HBAR)$0.0965425.71%
  • suiSui(SUI)$1.00-4.35%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.43-0.29%
  • BittensorBittensor(TAO)$315.0710.70%
  • shiba-inuShiba Inu(SHIB)$0.0000063.30%
  • crypto-com-chainCronos(CRO)$0.0662503.46%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • MemeCoreMemeCore(M)$1.31-12.87%
  • paypal-usdPayPal USD(PYUSD)$1.000.00%
  • tether-goldTether Gold(XAUT)$4,324.12-0.31%
  • okbOKB(OKB)$121.05-1.41%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • Ripple USDRipple USD(RLUSD)$1.000.01%
  • BitwayBitway(BTW)$0.88-2.67%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.14-0.18%
  • aaveAave(AAVE)$141.83-2.51%
  • mantleMantle(MNT)$0.653.28%
  • EthenaEthena(ENA)$0.206062-7.77%
  • OndoOndo(ONDO)$0.425605-5.15%
  • Pump.funPump.fun(PUMP)$0.0043820.02%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Enhancing Trust in Large Language Models: Fine-Tuning for Calibrated Uncertainties in High-Stakes Applications

June 16, 2024
in AI & Technology
Reading Time: 5 mins read
A A
Enhancing Trust in Large Language Models: Fine-Tuning for Calibrated Uncertainties in High-Stakes Applications
ShareShareShareShareShare

Large language models (LLMs) face a significant challenge in accurately representing uncertainty over the correctness of their output. This issue is critical for decision-making applications, particularly in fields like healthcare where erroneous confidence can lead to dangerous outcomes. The task is further complicated by linguistic variances in freeform generation, which cannot be exhaustively accounted for during training. LLM practitioners must navigate the dichotomy between black-box and white-box estimation methods, with the former gaining popularity due to restricted models, while the latter becoming more accessible with open-source models.

Existing attempts to address this challenge explored various approaches. Some methods utilize LLMs’ natural expression of distribution over possible outcomes, using predicted token probabilities for multiple-choice tests. However, these become less reliable for sentence-length answers due to the need to spread probabilities over many phrasings. Other approaches utilize prompting to produce uncertainty estimates, capitalizing on LLMs’ learned concepts of “correctness” and probabilities. Linear probes have also been used to classify a model’s correctness based on hidden representations. Despite these efforts, black-box methods often fail to generate useful uncertainties for popular open-source models, necessitating careful fine-tuning interventions.

To advance the debate on necessary interventions for good calibration, researchers from New York University, Abacus AI, and Cambridge University have conducted a deep investigation into the uncertainty calibration of LLMs. They propose fine-tuning for better uncertainties, which provides faster and more reliable estimates while using relatively few additional parameters. This method shows promise in generalizing to new question types and tasks beyond the fine-tuning dataset. The approach involves teaching language models to recognize what they don’t know using a calibration dataset, exploring effective parameterization, and determining the amount of data required for good generalization.

The proposed method involves focusing on black-box techniques for estimating a language model’s uncertainty, particularly those requiring a single sample or forward pass. For an open-ended generation, where answers are not limited to individual tokens or prescribed possibilities, researchers use perplexity as a length-normalized metric. The approach also explores prompting methods as an alternative to sequence likelihood, introducing formats that lay the foundation for recent work. These include zero-shot classifiers and verbalized confidence statements, which are used to create uncertainty estimates from language model outputs.

Results show that fine-tuning for uncertainties significantly improves performance compared to commonly used baselines. The quality of black-box uncertainty estimates produced by open-source models was examined against accuracy, using models like LLaMA-2, Mistral, and LLaMA-3. Evaluation on open-ended MMLU revealed that prompting methods typically give poorly calibrated uncertainties, with calibration not improving out-of-the-box as the base model improves. However, AUROC showed slight improvement with the power of the underlying model, although still lagging behind models with fine-tuning for uncertainty.

This study finds that out-of-the-box uncertainties from LLMs are unreliable for open-ended generation, contrary to prior results. The introduced fine-tuning procedures produce calibrated uncertainties with practical generalization properties. Notably, fine-tuning proves to be surprisingly sample-efficient and doesn’t rely on representations specific to a model evaluating its generations. The research also demonstrates the possibility of calibrated uncertainties being robust to distribution shifts. 


Check out the Paper. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter. 

Join our Telegram Channel and LinkedIn Group.

If you like our work, you will love our newsletter..

Don’t Forget to join our 44k+ ML SubReddit


YOU MAY ALSO LIKE

How To Enter VR Mode On Steam

Peloton Has Made A Foldable (Treadmill)

Asjad is an intern consultant at Marktechpost. He is persuing B.Tech in mechanical engineering at the Indian Institute of Technology, Kharagpur. Asjad is a Machine learning and deep learning enthusiast who is always researching the applications of machine learning in healthcare.


🐝 Join the Fastest Growing AI Research Newsletter Read by Researchers from Google + NVIDIA + Meta + Stanford + MIT + Microsoft and many others…


Credit: Source link

ShareTweetSendSharePin

Related Posts

How To Enter VR Mode On Steam
AI & Technology

How To Enter VR Mode On Steam

September 22, 2026
Peloton Has Made A Foldable (Treadmill)
AI & Technology

Peloton Has Made A Foldable (Treadmill)

September 22, 2026
OpenAI Faces Lawsuit From British Columbia Over Tumbler Ridge Shooting
AI & Technology

OpenAI Faces Lawsuit From British Columbia Over Tumbler Ridge Shooting

September 22, 2026
NVIDIA Introduces SoL-Pi: Auto-Research Loops That Cut Coding Agent Token Traffic by Up to 49%
AI & Technology

NVIDIA Introduces SoL-Pi: Auto-Research Loops That Cut Coding Agent Token Traffic by Up to 49%

September 22, 2026
Next Post
Speaker Johnson addresses meeting with Ukrainian President Zelenskyy

Speaker Johnson addresses meeting with Ukrainian President Zelenskyy

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Still The Best (And It’s Not Close)

Still The Best (And It’s Not Close)

September 18, 2026
What’s next for prosecutors after mistrial

What’s next for prosecutors after mistrial

September 17, 2026
Ted Cruz says ‘there will be a time’ to make a decision on a 2028 run for president

Ted Cruz says ‘there will be a time’ to make a decision on a 2028 run for president

September 21, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!