• bitcoinBitcoin(BTC)$84,414.000.66%
  • ethereumEthereum(ETH)$2,684.291.06%
  • tetherTether(USDT)$1.00-0.01%
  • binancecoinBNB(BNB)$779.962.15%
  • rippleXRP(XRP)$1.532.85%
  • usd-coinUSDC(USDC)$1.000.00%
  • solanaSolana(SOL)$117.032.69%
  • tronTRON(TRX)$0.3413980.75%
  • zcashZcash(ZEC)$1,546.301.17%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.02-1.27%
  • HyperliquidHyperliquid(HYPE)$93.611.56%
  • dogecoinDogecoin(DOGE)$0.0959254.38%
  • moneroMonero(XMR)$554.290.93%
  • whitebitWhiteBIT Coin(WBT)$84.510.32%
  • USDSUSDS(USDS)$1.000.02%
  • chainlinkChainlink(LINK)$12.895.61%
  • cardanoCardano(ADA)$0.2479864.57%
  • RainRain(RAIN)$0.012081-2.80%
  • leo-tokenLEO Token(LEO)$8.89-0.91%
  • stellarStellar(XLM)$0.2114274.37%
  • bitcoin-cashBitcoin Cash(BCH)$337.46-3.53%
  • nearNEAR Protocol(NEAR)$4.709.90%
  • uniswapUniswap(UNI)$9.261.88%
  • litecoinLitecoin(LTC)$73.9923.42%
  • Ethena USDeEthena USDe(USDE)$1.000.01%
  • avalanche-2Avalanche(AVAX)$10.421.74%
  • daiDai(DAI)$1.000.03%
  • CantonCanton(CC)$0.1125134.74%
  • USD1USD1(USD1)$1.000.00%
  • suiSui(SUI)$1.016.10%
  • hedera-hashgraphHedera(HBAR)$0.0925292.90%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.442.20%
  • shiba-inuShiba Inu(SHIB)$0.0000063.93%
  • BittensorBittensor(TAO)$292.030.41%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • crypto-com-chainCronos(CRO)$0.0627802.30%
  • BitwayBitway(BTW)$1.047.06%
  • MemeCoreMemeCore(M)$1.231.76%
  • paypal-usdPayPal USD(PYUSD)$1.000.01%
  • tether-goldTether Gold(XAUT)$4,272.72-0.21%
  • OndoOndo(ONDO)$0.5225.33%
  • okbOKB(OKB)$119.811.28%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • Ripple USDRipple USD(RLUSD)$1.000.02%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.14-0.20%
  • aaveAave(AAVE)$144.374.09%
  • mantleMantle(MNT)$0.673.95%
  • EthenaEthena(ENA)$0.2190095.53%
  • MorphoMorpho(MORPHO)$2.8511.12%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Apple Researchers Present KGLens: A Novel AI Method Tailored for Visualizing and Evaluating the Factual Knowledge Embedded in LLMs

August 12, 2024
in AI & Technology
Reading Time: 5 mins read
A A
Apple Researchers Present KGLens: A Novel AI Method Tailored for Visualizing and Evaluating the Factual Knowledge Embedded in LLMs
ShareShareShareShareShare

Large Language Models (LLMs) have gained significant attention for their versatility, but their factualness remains a critical concern. Studies have revealed that LLMs can produce nonfactual, hallucinated, or outdated information, undermining reliability. Current evaluation methods, such as fact-checking and fact-QA, face several challenges. Fact-checking struggles to assess the factualness of generated content, while fact-QA encounters difficulties scaling up evaluation data due to expensive annotation processes. Both approaches also face the risk of data contamination from web-crawled pretraining corpora. Also, LLMs often respond inconsistently to the same fact when presented in different forms, a challenge is that existing evaluation datasets need to be equipped to address.

Existing attempts to evaluate LLMs’ knowledge primarily use specific datasets, but face challenges like data leakage, static content, and limited metrics. Knowledge graphs (KGs) offer advantages in customization, evolving knowledge, and reduced test set leakage. Methods like LAMA and LPAQA use KGs for evaluation but struggle with unnatural question formats and impracticality for large KGs. KaRR overcomes some issues but remains inefficient for large graphs and lacks generalizability. Current approaches focus on accuracy over reliability, failing to address LLMs’ inconsistent responses to the same fact. Also, no existing work visualizes LLMs’ knowledge using KGs, presenting an opportunity for improvement. These limitations highlight the need for more comprehensive and efficient methods to evaluate and understand LLMs’ knowledge retention and accuracy.

Researchers from Apple introduced KGLENS, an innovative knowledge probing framework that has been developed to measure knowledge alignment between KGs and LLMs and identify LLMs’ knowledge blind spots. The framework employs a Thompson sampling-inspired method with a parameterized knowledge graph (PKG) to probe LLMs efficiently. KGLENS features a graph-guided question generator that converts KGs into natural language using GPT-4, designing two types of questions (fact-checking and fact-QA) to reduce answer ambiguity. Human evaluation shows that 97.7% of generated questions are sensible to annotators.

KGLENS employs a unique approach to efficiently probe LLMs’ knowledge using a PKG and Thompson sampling-inspired method. The framework initializes a PKG where each edge is augmented with a beta distribution, indicating the LLM’s potential deficiency on that edge. It then samples edges based on their probability, generates questions from these edges, and examines the LLM through a question-answering task. The PKG is updated based on the results, and this process iterates until convergence. Also, This framework features a graph-guided question generator that converts KG edges into natural language questions using GPT-4. It creates two types of questions: Yes/No questions for judgment and Wh-questions for generation, with the question type controlled by the graph structure. Entity aliases are included to reduce ambiguity.

For answer verification, KGLENS instructs LLMs to generate specific response formats and employs GPT-4 to check the correctness of responses for Wh-questions. The framework’s efficiency is evaluated through various sampling methods, demonstrating its effectiveness in identifying LLMs’ knowledge blind spots across diverse topics and relationships.

KGLENS evaluation across various LLMs reveals that the GPT-4 family consistently outperforms other models. GPT-4, GPT-4o, and GPT-4-turbo show comparable performance, with GPT-4o being more cautious with personal information. A significant gap exists between GPT-3.5-turbo and GPT-4, with GPT-3.5-turbo sometimes performing worse than legacy LLMs due to its conservative approach. Legacy models like Babbage-002 and Davinci-002 show only slight improvement over random guessing, highlighting the progress in recent LLMs. The evaluation provides insights into different error types and model behaviors, demonstrating the varying capabilities of LLMs in handling diverse knowledge domains and difficulty levels.

KGLENS introduces an efficient method for evaluating factual knowledge in LLMs using a Thompson sampling-inspired approach with parameterized Knowledge Graphs. The framework outperforms existing methods in revealing knowledge blind spots and demonstrates adaptability across various domains. Human evaluation confirms its effectiveness, achieving 95.7% accuracy. KGLENS and its assessment of KGs will be made available to the research community, fostering collaboration. For businesses, this tool facilitates the development of more reliable AI systems, enhancing user experiences and improving model knowledge. KGLENS represents a significant advancement in creating more accurate and dependable AI applications.


Check out the Paper. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter and join our Telegram Channel and LinkedIn Group. If you like our work, you will love our newsletter..

Don’t Forget to join our 48k+ ML SubReddit

Find Upcoming AI Webinars here



Asjad is an intern consultant at Marktechpost. He is persuing B.Tech in mechanical engineering at the Indian Institute of Technology, Kharagpur. Asjad is a Machine learning and deep learning enthusiast who is always researching the applications of machine learning in healthcare.

YOU MAY ALSO LIKE

From Anthropic to Robots: AI’s Next Frontier

Morgan Stanley’s Jonas: Physical AI Could Multiply Global GDP


Credit: Source link

ShareTweetSendSharePin

Related Posts

From Anthropic to Robots: AI’s Next Frontier
AI & Technology

From Anthropic to Robots: AI’s Next Frontier

September 24, 2026
Morgan Stanley’s Jonas: Physical AI Could Multiply Global GDP
AI & Technology

Morgan Stanley’s Jonas: Physical AI Could Multiply Global GDP

September 24, 2026
AI Agents Fuel a New Cybersecurity Boom
AI & Technology

AI Agents Fuel a New Cybersecurity Boom

September 24, 2026
Bessemer: Anthropic Has Been Consistent on AI Safety
AI & Technology

Bessemer: Anthropic Has Been Consistent on AI Safety

September 24, 2026
Next Post
Goat named ‘Jeffrey’ trapped for hours on ledge in Kansas City

Goat named ‘Jeffrey’ trapped for hours on ledge in Kansas City

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Mamdani announces one-year ban on AI in public schools

Mamdani announces one-year ban on AI in public schools

September 19, 2026
Vanderbilt wins in stunning fashion after NC State QB CJ Bailey fumbles in end zone with 1 second left – Yahoo Sports

Vanderbilt wins in stunning fashion after NC State QB CJ Bailey fumbles in end zone with 1 second left – Yahoo Sports

September 19, 2026
Odd Lots: OpenAI’s Brockman Is Optimistic About AI Development

Odd Lots: OpenAI’s Brockman Is Optimistic About AI Development

September 20, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!