• bitcoinBitcoin(BTC)$80,450.00-1.02%
  • ethereumEthereum(ETH)$2,576.42-2.40%
  • tetherTether(USDT)$1.00-0.02%
  • binancecoinBNB(BNB)$752.21-2.02%
  • rippleXRP(XRP)$1.38-3.96%
  • usd-coinUSDC(USDC)$1.00-0.02%
  • solanaSolana(SOL)$108.49-3.03%
  • tronTRON(TRX)$0.3422721.33%
  • zcashZcash(ZEC)$1,444.48-6.28%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.02-1.32%
  • HyperliquidHyperliquid(HYPE)$91.40-0.94%
  • dogecoinDogecoin(DOGE)$0.085107-3.65%
  • moneroMonero(XMR)$523.72-10.95%
  • whitebitWhiteBIT Coin(WBT)$81.82-1.60%
  • USDSUSDS(USDS)$1.00-0.02%
  • RainRain(RAIN)$0.013184-5.38%
  • chainlinkChainlink(LINK)$12.06-3.78%
  • cardanoCardano(ADA)$0.220759-2.68%
  • leo-tokenLEO Token(LEO)$8.960.77%
  • stellarStellar(XLM)$0.190318-2.26%
  • uniswapUniswap(UNI)$8.77-3.76%
  • bitcoin-cashBitcoin Cash(BCH)$246.32-2.19%
  • Ethena USDeEthena USDe(USDE)$1.00-0.01%
  • nearNEAR Protocol(NEAR)$3.63-1.59%
  • daiDai(DAI)$1.00-0.01%
  • avalanche-2Avalanche(AVAX)$10.219.18%
  • litecoinLitecoin(LTC)$57.14-1.47%
  • USD1USD1(USD1)$1.00-0.04%
  • CantonCanton(CC)$0.104363-6.01%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.37-0.72%
  • hedera-hashgraphHedera(HBAR)$0.0811660.48%
  • suiSui(SUI)$0.82-4.74%
  • MemeCoreMemeCore(M)$1.4813.95%
  • Global DollarGlobal Dollar(USDG)$1.00-0.01%
  • shiba-inuShiba Inu(SHIB)$0.000005-2.23%
  • crypto-com-chainCronos(CRO)$0.058079-3.18%
  • BittensorBittensor(TAO)$251.12-6.92%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.02%
  • tether-goldTether Gold(XAUT)$4,368.46-0.11%
  • Circle USYCCircle USYC(USYC)$1.140.00%
  • okbOKB(OKB)$115.95-5.36%
  • Ripple USDRipple USD(RLUSD)$1.00-0.01%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.14-0.32%
  • EthenaEthena(ENA)$0.2081168.88%
  • aaveAave(AAVE)$135.74-5.43%
  • OndoOndo(ONDO)$0.4123680.55%
  • AsterAster(ASTER)$0.74-4.10%
  • mantleMantle(MNT)$0.59-3.08%
  • BitwayBitway(BTW)$0.7216.82%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

THRONE: Advancing the Evaluation of Hallucinations in Vision-Language Models

May 12, 2024
in AI & Technology
Reading Time: 5 mins read
A A
THRONE: Advancing the Evaluation of Hallucinations in Vision-Language Models
ShareShareShareShareShare

Understanding and mitigating hallucinations in vision-language models (VLVMs) is an emerging field of research that addresses the generation of coherent but factually incorrect responses by these advanced AI systems. As VLVMs increasingly integrate text and visual inputs to generate responses, the accuracy of these outputs becomes crucial, especially in settings where precision is paramount, such as medical diagnostics or autonomous driving.

Hallucinations in VLVMs typically manifest as plausible yet incorrect details generated about an image. These inaccuracies pose significant risks, potentially misinforming decisions in critical applications. The challenge lies in detecting these errors and developing methods to mitigate them effectively, ensuring the reliability of VLVM outputs.

Most existing benchmarks for evaluating hallucinations in VLVMs focus on responses to constrained query formats, such as yes/no questions about specific objects or attributes within an image. These benchmarks often fail to measure more complex, open-ended hallucinations that can occur in varied real-world applications. As a result, there is a significant gap in the ability to fully understand and mitigate the broader spectrum of hallucinations that VLVMs can produce.

Researchers from the University of Oxford, AWS AI Labs, introduced a new framework called THRONE (Text-from-image Hallucination Recognition with Object-probes for open-ended Evaluation) to address this gap. THRONE is designed to assess Type I hallucinations, those that occur in response to open-ended prompts requiring detailed image descriptions. Unlike previous methods, THRONE uses publicly available language models to evaluate the hallucinations in free-form responses generated by various VLVMs, offering a more comprehensive and rigorous approach.

THRONE leverages multiple metrics to measure hallucinations across different VLVMs quantitatively. For example, it employs precision and recall metrics alongside a class-wise F0.5 score, emphasizing precision twice as much as recall. This scoring is particularly relevant in scenarios where false positives, incorrect but plausible responses, are more detrimental than false negatives.

An evaluation of THRONE’s effectiveness revealed insightful data about the prevalence and characteristics of hallucinations in current VLVMs. Despite the framework’s advanced approach, the results indicate that many VLVMs still struggle with a high rate of hallucinations. For instance, the framework detected that some of the evaluated models produce responses, with about 20% of the objects mentioned being hallucinations. This high rate of inaccuracies underscores the persistent challenge of reducing hallucinations and improving the reliability of VLVM outputs.

In conclusion, the THRONE framework represents a significant step forward in evaluating hallucinations in vision-language models, particularly addressing the complex issue of Type I hallucinations in free-form responses. While existing benchmarks have struggled to effectively measure these more nuanced errors, THRONE utilizes a novel combination of publicly available language models and a robust metric system, including precision, recall, and class-wise F0.5 scores. Despite these advances, the high rate of detected hallucinations, around 20% in some models, underscores the ongoing challenges and the necessity for further research to enhance the accuracy and reliability of VLVMs in practical applications.


Check out the Paper. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter. Join our Telegram Channel, Discord Channel, and LinkedIn Group.

If you like our work, you will love our newsletter..

Don’t Forget to join our 42k+ ML SubReddit


YOU MAY ALSO LIKE

How China Is Changing AI

Snap Makes Its Case for Wearing A Computer On Your Face

Sana Hassan, a consulting intern at Marktechpost and dual-degree student at IIT Madras, is passionate about applying technology and AI to address real-world challenges. With a keen interest in solving practical problems, he brings a fresh perspective to the intersection of AI and real-life solutions.


[Recommended Read] Rightsify’s GCX: Your Go-To Source for High-Quality, Ethically Sourced, Copyright-Cleared AI Music Training Datasets with Rich Metadata


Credit: Source link

ShareTweetSendSharePin

Related Posts

How China Is Changing AI
AI & Technology

How China Is Changing AI

September 20, 2026
Snap Makes Its Case for Wearing A Computer On Your Face
AI & Technology

Snap Makes Its Case for Wearing A Computer On Your Face

September 20, 2026
Trump Opposes AI Guardrails Amid Chip Selloff
AI & Technology

Trump Opposes AI Guardrails Amid Chip Selloff

September 20, 2026
Anthropic’s Claude Takes Bigger Role in Building AI
AI & Technology

Anthropic’s Claude Takes Bigger Role in Building AI

September 20, 2026
Next Post
Top Democratic lawmaker: Israel ‘not sufficiently concerned’ about Palestinian people right now

Top Democratic lawmaker: Israel ‘not sufficiently concerned’ about Palestinian people right now

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
More Iran War Strikes, Higher Oil Prices; Why an Algorithm is Cutting Disability Benefits | Sept. 9

More Iran War Strikes, Higher Oil Prices; Why an Algorithm is Cutting Disability Benefits | Sept. 9

September 14, 2026
Ingram Micro Holding Corporation (INGM) Analyst/Investor Day – Slideshow

Ingram Micro Holding Corporation (INGM) Analyst/Investor Day – Slideshow

September 18, 2026
A Princeton Researcher Proposes Recurrent Looped Transformer (RLT) that Carries Decoder State across Every Token, Fixing 96 Blocks per Token with Unbounded Temporal Depth

A Princeton Researcher Proposes Recurrent Looped Transformer (RLT) that Carries Decoder State across Every Token, Fixing 96 Blocks per Token with Unbounded Temporal Depth

September 13, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!