• bitcoinBitcoin(BTC)$85,953.000.70%
  • ethereumEthereum(ETH)$2,752.590.62%
  • tetherTether(USDT)$1.000.00%
  • binancecoinBNB(BNB)$789.72-0.06%
  • rippleXRP(XRP)$1.543.61%
  • usd-coinUSDC(USDC)$1.000.01%
  • solanaSolana(SOL)$117.370.46%
  • tronTRON(TRX)$0.3456220.33%
  • zcashZcash(ZEC)$1,528.61-0.68%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.011.27%
  • HyperliquidHyperliquid(HYPE)$95.770.40%
  • dogecoinDogecoin(DOGE)$0.0992836.35%
  • moneroMonero(XMR)$580.381.68%
  • whitebitWhiteBIT Coin(WBT)$86.500.00%
  • chainlinkChainlink(LINK)$12.97-1.13%
  • USDSUSDS(USDS)$1.000.00%
  • RainRain(RAIN)$0.013472-4.54%
  • cardanoCardano(ADA)$0.2482131.01%
  • leo-tokenLEO Token(LEO)$8.980.73%
  • stellarStellar(XLM)$0.2115701.03%
  • bitcoin-cashBitcoin Cash(BCH)$301.4111.15%
  • nearNEAR Protocol(NEAR)$4.589.87%
  • uniswapUniswap(UNI)$9.557.01%
  • Ethena USDeEthena USDe(USDE)$1.00-0.03%
  • avalanche-2Avalanche(AVAX)$10.83-5.35%
  • litecoinLitecoin(LTC)$60.83-3.17%
  • CantonCanton(CC)$0.1184902.49%
  • daiDai(DAI)$1.000.00%
  • USD1USD1(USD1)$1.00-0.03%
  • suiSui(SUI)$1.02-1.39%
  • hedera-hashgraphHedera(HBAR)$0.0947203.53%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.430.26%
  • BittensorBittensor(TAO)$321.3911.46%
  • shiba-inuShiba Inu(SHIB)$0.0000065.28%
  • crypto-com-chainCronos(CRO)$0.0664624.38%
  • Global DollarGlobal Dollar(USDG)$1.000.01%
  • MemeCoreMemeCore(M)$1.32-11.94%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.01%
  • tether-goldTether Gold(XAUT)$4,338.61-0.53%
  • okbOKB(OKB)$122.390.30%
  • Circle USYCCircle USYC(USYC)$1.140.02%
  • Ripple USDRipple USD(RLUSD)$1.000.01%
  • BitwayBitway(BTW)$0.875.79%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.15-0.28%
  • aaveAave(AAVE)$143.68-2.75%
  • mantleMantle(MNT)$0.664.13%
  • EthenaEthena(ENA)$0.212088-6.95%
  • Pump.funPump.fun(PUMP)$0.0045654.28%
  • OndoOndo(ONDO)$0.432449-4.81%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Understanding Hallucination Rates in Language Models: Insights from Training on Knowledge Graphs and Their Detectability Challenges

August 18, 2024
in AI & Technology
Reading Time: 5 mins read
A A
Understanding Hallucination Rates in Language Models: Insights from Training on Knowledge Graphs and Their Detectability Challenges
ShareShareShareShareShare

Language models (LMs) exhibit improved performance with increased size and training data, yet the relationship between model scale and hallucinations remains unexplored. Defining hallucinations in LMs presents challenges due to their varied manifestations. A new study from Google Deepmind focuses on hallucinations where correct answers appear verbatim in training data. Achieving low hallucination rates demands larger models and more computational resources than previously thought. Hallucination detection becomes increasingly difficult as LM size grows. Knowledge graphs (KGs) offer a promising approach to providing structured, factual training data for LMs, potentially mitigating hallucinations.

The study investigates the relationship between the language model (LM) scale and hallucinations, focusing on instances where correct answers are present in the training data. Using a knowledge graph (KG)–based dataset, researchers train increasingly large LMs to control training content effectively. Findings indicate that larger, longer-trained LMs hallucinate less, but achieving low hallucination rates requires significantly more resources than previously thought. The study also reveals an inverse relationship between the LM scale and hallucination detectability.

YOU MAY ALSO LIKE

Peloton Has Made A Foldable (Treadmill)

OpenAI Faces Lawsuit From British Columbia Over Tumbler Ridge Shooting

Precisely defining and quantifying hallucinations in natural language settings remains challenging due to language ambiguity and unclear knowledge content in training data. Despite advancements in generative capabilities, hallucinations persist as a significant challenge for LMs. The research addresses the gap in understanding how hallucinations depend on model scale. Knowledge graphs offer a structured approach to LM training, enabling straightforward fact verification against the dataset and providing a quantifiable measure of hallucination.

Traditional language models (LMs) trained on natural language data often produce hallucinations and repetitive information due to semantic ambiguity. The study employs a knowledge graph (KG) approach, using structured triplets of information to provide a clearer understanding of how LMs misrepresent training data. This method allows for a more precise evaluation of hallucinations and their relationship to model scale.

The study constructs a dataset using knowledge graph triplets (subject, predicate, object), enabling precise control over training data and quantifiable hallucination measurement. Language models (LMs) are trained from scratch on this dataset, optimizing auto-regressive log-likelihood. Evaluation involves prompting models with subject and predicate, and assessing object completion accuracy against the knowledge graph. Token tasks and head detectors evaluate hallucination detection performance. The methodology focuses on hallucinations where correct answers appear verbatim in the training set, exploring the relationship between the LM scale and hallucination frequency.

The research trains increasingly large LMs to investigate scale effects on hallucination rates and detectability. Analysis reveals that larger, longer-trained LMs hallucinate less, though larger datasets may increase hallucination rates. The authors acknowledge limitations in generalizability to all hallucination types and the use of smaller-than-state-of-the-art models. This comprehensive approach provides insights into LM hallucinations and their detectability, contributing to the field of natural language processing.

The study reveals that larger language models and extended training reduce hallucinations on fixed datasets, while increased dataset size elevates hallucination rates. Hallucination detectors show high accuracy, improving with model size. Token-level detection generally outperforms other methods. A trade-off exists between fact recall and generalization ability, with extended training minimizing hallucinations on seen data but risking overfitting on unseen data. AUC-PR serves as a reliable measure of detector performance. These findings highlight the complex relationship between model scale, dataset size, and hallucination rates, emphasizing the importance of balancing model size and training duration to mitigate hallucinations while addressing challenges posed by larger datasets.

In conclusion, the study reveals that larger, longer-trained language models exhibit reduced hallucination rates, but achieving minimal hallucinations requires substantial computational resources. Increased dataset size correlates with higher hallucination rates when model size and training epochs remain constant. A trade-off exists between memorization and generalization, with extended training improving fact retention but potentially hindering adaptability to new data. Paradoxically, as models grow larger and hallucinate less, detecting remaining hallucinations becomes more challenging. Future research should focus on enhancing hallucination detection in larger models and exploring the practical implications of these findings for language model applications.


Check out the Paper. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter and join our Telegram Channel and LinkedIn Group. If you like our work, you will love our newsletter..

Don’t Forget to join our 48k+ ML SubReddit

Find Upcoming AI Webinars here



Shoaib Nazir is a consulting intern at MarktechPost and has completed his M.Tech dual degree from the Indian Institute of Technology (IIT), Kharagpur. With a strong passion for Data Science, he is particularly interested in the diverse applications of artificial intelligence across various domains. Shoaib is driven by a desire to explore the latest technological advancements and their practical implications in everyday life. His enthusiasm for innovation and real-world problem-solving fuels his continuous learning and contribution to the field of AI


Credit: Source link

ShareTweetSendSharePin

Related Posts

Peloton Has Made A Foldable (Treadmill)
AI & Technology

Peloton Has Made A Foldable (Treadmill)

September 22, 2026
OpenAI Faces Lawsuit From British Columbia Over Tumbler Ridge Shooting
AI & Technology

OpenAI Faces Lawsuit From British Columbia Over Tumbler Ridge Shooting

September 22, 2026
NVIDIA Introduces SoL-Pi: Auto-Research Loops That Cut Coding Agent Token Traffic by Up to 49%
AI & Technology

NVIDIA Introduces SoL-Pi: Auto-Research Loops That Cut Coding Agent Token Traffic by Up to 49%

September 22, 2026
SpaceXAI Releases Grok 4.7: A Larger Base Model at the Same / Price as Grok 4.6
AI & Technology

SpaceXAI Releases Grok 4.7: A Larger Base Model at the Same $2/$6 Price as Grok 4.6

September 22, 2026
Next Post
EmBARDiment: An Implicit Attention Framework that Enhances AI Interaction Efficiency in Extended Reality Through Eye-Tracking and Contextual Memory Integration

EmBARDiment: An Implicit Attention Framework that Enhances AI Interaction Efficiency in Extended Reality Through Eye-Tracking and Contextual Memory Integration

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Standalone AR Glasses Are Here

Standalone AR Glasses Are Here

September 16, 2026
Stanford Researchers Release Paper2Agent: Turning Research Papers Into AI Agents That Reproduce Results and Run on New Data

Stanford Researchers Release Paper2Agent: Turning Research Papers Into AI Agents That Reproduce Results and Run on New Data

September 16, 2026
Powerful waves from Hurricane Marie slam Southern California

Powerful waves from Hurricane Marie slam Southern California

September 16, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!