• bitcoinBitcoin(BTC)$85,479.004.87%
  • ethereumEthereum(ETH)$2,733.772.77%
  • tetherTether(USDT)$1.000.00%
  • binancecoinBNB(BNB)$787.311.97%
  • rippleXRP(XRP)$1.526.45%
  • usd-coinUSDC(USDC)$1.000.00%
  • solanaSolana(SOL)$116.654.29%
  • tronTRON(TRX)$0.3489641.72%
  • zcashZcash(ZEC)$1,503.65-0.18%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.011.29%
  • HyperliquidHyperliquid(HYPE)$94.320.53%
  • dogecoinDogecoin(DOGE)$0.10005112.64%
  • moneroMonero(XMR)$577.96-8.19%
  • whitebitWhiteBIT Coin(WBT)$85.953.25%
  • RainRain(RAIN)$0.013673-2.42%
  • chainlinkChainlink(LINK)$12.953.54%
  • USDSUSDS(USDS)$1.00-0.01%
  • cardanoCardano(ADA)$0.2460665.76%
  • leo-tokenLEO Token(LEO)$8.950.41%
  • stellarStellar(XLM)$0.2125807.15%
  • nearNEAR Protocol(NEAR)$4.403.02%
  • uniswapUniswap(UNI)$8.893.00%
  • bitcoin-cashBitcoin Cash(BCH)$265.584.71%
  • Ethena USDeEthena USDe(USDE)$1.00-0.04%
  • avalanche-2Avalanche(AVAX)$10.74-3.62%
  • litecoinLitecoin(LTC)$60.994.41%
  • CantonCanton(CC)$0.1185414.99%
  • daiDai(DAI)$1.000.01%
  • USD1USD1(USD1)$1.00-0.05%
  • suiSui(SUI)$1.026.07%
  • hedera-hashgraphHedera(HBAR)$0.0930217.03%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.442.61%
  • BittensorBittensor(TAO)$320.7018.99%
  • shiba-inuShiba Inu(SHIB)$0.0000068.58%
  • crypto-com-chainCronos(CRO)$0.0660096.87%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • MemeCoreMemeCore(M)$1.37-9.49%
  • paypal-usdPayPal USD(PYUSD)$1.000.00%
  • tether-goldTether Gold(XAUT)$4,321.51-0.75%
  • okbOKB(OKB)$121.351.61%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • Ripple USDRipple USD(RLUSD)$1.00-0.01%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.15-0.47%
  • aaveAave(AAVE)$142.382.36%
  • BitwayBitway(BTW)$0.8014.03%
  • pepePepe(PEPE)$0.00000526.83%
  • EthenaEthena(ENA)$0.212285-1.62%
  • mantleMantle(MNT)$0.645.58%
  • OndoOndo(ONDO)$0.4337160.59%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Enhancing LLM Reliability: Detecting Confabulations with Semantic Entropy

June 22, 2024
in AI & Technology
Reading Time: 5 mins read
A A
Enhancing LLM Reliability: Detecting Confabulations with Semantic Entropy
ShareShareShareShareShare

LLMs like ChatGPT and Gemini demonstrate impressive reasoning and answering capabilities but often produce “hallucinations,” meaning they generate false or unsupported information. This problem hampers their reliability in critical fields, from law to medicine, where inaccuracies can have severe consequences. Efforts to reduce these errors through supervision or reinforcement have seen limited success. A subset of hallucinations, termed “confabulations,” involves LLMs giving arbitrary or incorrect responses to identical queries, such as varying answers to a medical question about Sotorasib. This issue is distinct from errors caused by training on faulty data or systematic reasoning failures. Understanding and addressing these nuanced error types is crucial for improving LLM reliability.

Researchers from the OATML group at the University of Oxford have developed a statistical approach to detect a specific type of error in LLMs, known as “confabulations.” These errors occur when LLMs generate arbitrary and incorrect responses, often due to subtle variations in the input or random seed. The new method leverages entropy-based uncertainty estimators, focusing on the meaning rather than the exact wording of responses. By assessing the “semantic entropy” — the uncertainty in the sense of generated answers — this technique can identify when LLMs are likely to produce unreliable outputs. This method does not require knowledge of the specific task or labeled data and is effective across different datasets and applications. It improves LLM reliability by signaling when extra caution is needed, thus allowing users to avoid or critically evaluate potentially confabulated answers.

YOU MAY ALSO LIKE

SpaceXAI Releases Grok 4.7: A Larger Base Model at the Same $2/$6 Price as Grok 4.6

Why It’s Important To Unplug Your PC During A Power Outage

The researchers’ method works by clustering similar answers based on their meaning and measuring the entropy within these clusters. If the entropy is high, the LLM is likely generating confabulated responses. This process enhances the detection of semantic inconsistencies that naive entropy measures, which only consider lexical differences, might miss. The technique has been tested on various LLMs across multiple domains, such as trivia, general knowledge, and medical queries, demonstrating significant improvements in detecting and filtering unreliable answers. Moreover, by refusing to answer questions likely to produce high-entropy (confabulated) responses, the method can enhance the overall accuracy of LLM outputs. This innovation represents a critical advancement in ensuring the reliability of LLMs, particularly in free-form text generation where traditional supervised learning methods fall short.

Semantic entropy is a method to detect confabulations in LLMs by measuring their uncertainty over the meaning of generated outputs. This technique leverages predictive entropy and clusters generated sequences by semantic equivalence using bidirectional entailment. It computes semantic entropy based on the probabilities of these clusters, indicating the model’s confidence in its answers. By sampling outputs and clustering them, semantic entropy identifies when a model’s answers are likely arbitrary. This approach helps predict model accuracy, improves reliability by flagging uncertain answers, and gives users a better confidence assessment of model outputs.

The study focuses on identifying and mitigating confabulations—erroneous or misleading outputs—generated by LLMs using a metric called “semantic entropy.” This metric evaluates the variability in meaning across different generations of model outputs, distinguishing it from traditional entropy measures that only consider lexical differences. The research shows that semantic entropy, which accounts for consistent meaning despite diverse phrasings, effectively detects when LLMs produce incorrect or misleading responses. Semantic entropy outperformed baseline methods like naive entropy and supervised embedding regression across various datasets and model sizes, including LLaMA, Falcon, and Mistral models, outperforming baseline methods like naive entropy and supervised embedding regression, achieving a notable AUROC 0.790. This suggests that semantic entropy provides a robust mechanism for identifying confabulations, even in distribution shifts between training and deployment.

Moreover, the study extends the application of semantic entropy to longer text passages, such as biographical paragraphs, by breaking them into factual claims and evaluating the consistency of these claims through rephrasing. This approach demonstrated that semantic entropy could effectively detect confabulations in extended text, outperforming simple self-check mechanisms and adapting probability-based methods. The findings imply that LLMs inherently possess the ability to recognize their knowledge gaps, but traditional evaluation methods may only partially leverage this capacity. Thus, semantic entropy offers a promising direction for improving the reliability of LLM outputs in complex and open-ended tasks, providing a way to assess and manage the uncertainties in their responses.


Check out the Paper, Project, and GitHub. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter. 

Join our Telegram Channel and LinkedIn Group.

If you like our work, you will love our newsletter..

Don’t Forget to join our 45k+ ML SubReddit


Sana Hassan, a consulting intern at Marktechpost and dual-degree student at IIT Madras, is passionate about applying technology and AI to address real-world challenges. With a keen interest in solving practical problems, he brings a fresh perspective to the intersection of AI and real-life solutions.

🐝 Join the Fastest Growing AI Research Newsletter Read by Researchers from Google + NVIDIA + Meta + Stanford + MIT + Microsoft and many others…


Credit: Source link

ShareTweetSendSharePin

Related Posts

SpaceXAI Releases Grok 4.7: A Larger Base Model at the Same / Price as Grok 4.6
AI & Technology

SpaceXAI Releases Grok 4.7: A Larger Base Model at the Same $2/$6 Price as Grok 4.6

September 22, 2026
Why It’s Important To Unplug Your PC During A Power Outage
AI & Technology

Why It’s Important To Unplug Your PC During A Power Outage

September 22, 2026
Why Is Your Laptop Fan So Loud?
AI & Technology

Why Is Your Laptop Fan So Loud?

September 22, 2026
These Drones Could Cover Up To 98 Percent Of The World’s Oceans By 2028
AI & Technology

These Drones Could Cover Up To 98 Percent Of The World’s Oceans By 2028

September 21, 2026
Next Post
Internal probe clears Ohio officers who fatally shot Jayland Walker

Internal probe clears Ohio officers who fatally shot Jayland Walker

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Stock Market Today: Dow Steady, Yields Edge Higher — Live Updates – WSJ

Stock Market Today: Dow Steady, Yields Edge Higher — Live Updates – WSJ

September 18, 2026
Dolly Parton-themed corn maze honors singer after her death

Dolly Parton-themed corn maze honors singer after her death

September 21, 2026
Plane with an Amazon logo crashes outside Miami airport

Plane with an Amazon logo crashes outside Miami airport

September 16, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!