• bitcoinBitcoin(BTC)$83,203.000.19%
  • ethereumEthereum(ETH)$2,668.200.22%
  • tetherTether(USDT)$1.00-0.01%
  • binancecoinBNB(BNB)$758.520.07%
  • rippleXRP(XRP)$1.490.36%
  • usd-coinUSDC(USDC)$1.00-0.01%
  • solanaSolana(SOL)$119.021.20%
  • tronTRON(TRX)$0.3351250.28%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.032.76%
  • zcashZcash(ZEC)$1,409.021.99%
  • HyperliquidHyperliquid(HYPE)$85.95-0.62%
  • dogecoinDogecoin(DOGE)$0.0932100.29%
  • chainlinkChainlink(LINK)$14.36-3.99%
  • moneroMonero(XMR)$541.170.05%
  • USDSUSDS(USDS)$1.000.00%
  • whitebitWhiteBIT Coin(WBT)$83.190.25%
  • cardanoCardano(ADA)$0.2435650.53%
  • RainRain(RAIN)$0.0125811.41%
  • leo-tokenLEO Token(LEO)$9.030.00%
  • stellarStellar(XLM)$0.219848-2.59%
  • nearNEAR Protocol(NEAR)$4.946.92%
  • bitcoin-cashBitcoin Cash(BCH)$305.39-0.02%
  • uniswapUniswap(UNI)$8.741.63%
  • litecoinLitecoin(LTC)$66.92-1.23%
  • CantonCanton(CC)$0.127481-4.28%
  • avalanche-2Avalanche(AVAX)$11.257.89%
  • Ethena USDeEthena USDe(USDE)$1.000.00%
  • suiSui(SUI)$1.153.11%
  • daiDai(DAI)$1.000.00%
  • hedera-hashgraphHedera(HBAR)$0.103305-13.20%
  • USD1USD1(USD1)$1.00-0.01%
  • quant-networkQuant(QNT)$294.1629.07%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.48-3.93%
  • BitwayBitway(BTW)$1.4023.52%
  • BittensorBittensor(TAO)$300.590.03%
  • shiba-inuShiba Inu(SHIB)$0.0000062.80%
  • tether-goldTether Gold(XAUT)$4,178.870.94%
  • crypto-com-chainCronos(CRO)$0.067062-1.19%
  • Global DollarGlobal Dollar(USDG)$1.00-0.01%
  • paypal-usdPayPal USD(PYUSD)$1.000.00%
  • Pump.funPump.fun(PUMP)$0.00579622.02%
  • okbOKB(OKB)$120.802.42%
  • Ripple USDRipple USD(RLUSD)$1.00-0.01%
  • EthenaEthena(ENA)$0.245352-1.59%
  • aaveAave(AAVE)$159.567.92%
  • OndoOndo(ONDO)$0.497938-1.14%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • MemeCoreMemeCore(M)$1.04-5.64%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.15-0.06%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Will LLMs Replace Knowledge Graphs? Meta Researchers Propose ‘Head-to-Tail’: A New Benchmark to Measure the Factual Knowledge of Large Language Models

August 30, 2023
in AI & Technology
Reading Time: 4 mins read
A A
Will LLMs Replace Knowledge Graphs? Meta Researchers Propose ‘Head-to-Tail’: A New Benchmark to Measure the Factual Knowledge of Large Language Models
ShareShareShareShareShare

Large Language Models have gathered a lot of appreciation for their super amazing capabilities. They are able to imitate humans and generate content just like a human would do. Pre-trained large language models (LLMs), such as ChatGPT and LLaMA, have demonstrated astounding aptitudes for understanding the material and responding to frequent queries. Several studies have demonstrated their aptitude for internalizing knowledge and responding to inquiries. Though LLMs have significantly advanced, they frequently lack a sophisticated understanding of domain-specific nuances and are prone to producing incorrect information, known as hallucinations. This highlights the significant obstacles to improving LLM accuracy and reducing the incidence of hallucinating responses.

Discussion related to LLMs has majorly focused on three main areas, which are reducing hallucinations in LLM-generated responses, improving the factual accuracy of LLMs, and speculating on whether LLMs might eventually replace Knowledge Graphs (KGs) as a means of storing world knowledge in a symbolic format. Recently, a team of researchers from Meta Reality Labs have opted for a fresh approach to answer these questions by attempting to determine how much information LLMs actually possess.

While answering the question of how well-versed LLMs are in terms of knowledge, the team has discussed two aspects. Firstly, it can be difficult to directly question the knowledge contained within an LLM at first. Even if the knowledge is already incorporated in the model’s parameters, hallucinations could be caused by a lack of knowledge or a malfunctioning generative model. The study suggests using correctness as a metric to roughly gauge the degree of knowledge within an LLM. This involves assessing the model’s ability to answer clear, accurate questions like “Where was basketball player Michael Jordan born?” The LLM is also asked to provide succinct responses and admit uncertainty by using the word ‘unsure’ when its confidence is low.

Secondly, there is no readily accessible benchmark that accurately reflects the diversity of user interests or the breadth of information in the world. Even the most comprehensive knowledge graphs show gaps in knowledge, particularly when it comes to less well-known facts. The query logs from major LLMs or search engines are not publicly available.

To address all the limitations, the team has introduced a benchmark they have created called “Head-to-Tail.” This benchmark consists of a collection of 18,000 question-answer (QA) pairs that have been divided into head, torso, and tail facts based on the popularity of their respective subjects. Different public familiarity levels are reflected in these categories. The team has created an automated evaluation method and a set of measures that closely reflect the breadth of knowledge that an LLM has competently assimilated in order to evaluate the knowledge maintained by LLMs.

The research’s core is the evaluation of 14 LLMs that are available to the general public. The results showed that existing LLMs still need to improve significantly in terms of perfecting their comprehension of factual data. This is especially true for information that falls within the torso-to-tail area and concerns less well-known organizations.

In conclusion, this research examines the factual knowledge of LLMs using a recently proposed benchmark and cutting-edge evaluation techniques. The work makes a substantial contribution to the continuing discussion regarding the dependability and prospective advancements of big language models in incorporating factual information by addressing significant research problems and outlining specific findings.


Check out the Paper. All Credit For This Research Goes To the Researchers on This Project. Also, don’t forget to join our 29k+ ML SubReddit, 40k+ Facebook Community, Discord Channel, and Email Newsletter, where we share the latest AI research news, cool AI projects, and more.

If you like our work, you will love our newsletter..


YOU MAY ALSO LIKE

One Bad Prompt Took Down a Company’s Salesforce: RSA’s Jim Taylor on Agent ID and Taming the 4,000 Shadow AI Agents Hiding in Your Enterprise

Don’t Throw Away Your Old Router — Do This Instead

Tanya Malhotra is a final year undergrad from the University of Petroleum & Energy Studies, Dehradun, pursuing BTech in Computer Science Engineering with a specialization in Artificial Intelligence and Machine Learning.
She is a Data Science enthusiast with good analytical and critical thinking, along with an ardent interest in acquiring new skills, leading groups, and managing work in an organized manner.


🚀 CodiumAI enables busy developers to generate meaningful tests (Sponsored)

Credit: Source link

ShareTweetSendSharePin

Related Posts

One Bad Prompt Took Down a Company’s Salesforce: RSA’s Jim Taylor on Agent ID and Taming the 4,000 Shadow AI Agents Hiding in Your Enterprise
AI & Technology

One Bad Prompt Took Down a Company’s Salesforce: RSA’s Jim Taylor on Agent ID and Taming the 4,000 Shadow AI Agents Hiding in Your Enterprise

September 30, 2026
Don’t Throw Away Your Old Router — Do This Instead
AI & Technology

Don’t Throw Away Your Old Router — Do This Instead

September 30, 2026
The AI Industry Wants Models To Assist In Legal Battles, But Will They Help?
AI & Technology

The AI Industry Wants Models To Assist In Legal Battles, But Will They Help?

September 29, 2026
Liquid AI Releases d1: A Decision Model That Returns Calibrated Probabilities With Zero Output Tokens
AI & Technology

Liquid AI Releases d1: A Decision Model That Returns Calibrated Probabilities With Zero Output Tokens

September 29, 2026
Next Post
MTP75 Archives — Gloria Steinem: Women’s Rights Movement Not A ‘Militant Movement’

MTP75 Archives — Gloria Steinem: Women’s Rights Movement Not A ‘Militant Movement’

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
OpenAI Reportedly Cancels GPT-6.1 Astra’s Release Over Deceptive Behavior

OpenAI Reportedly Cancels GPT-6.1 Astra’s Release Over Deceptive Behavior

September 29, 2026
Anthropic Releases Claude Sonnet 5.5: 70.6% on Terminal-Bench 4.0 at the Same / Price

Anthropic Releases Claude Sonnet 5.5: 70.6% on Terminal-Bench 4.0 at the Same $2/$10 Price

September 29, 2026
Emerging Market Debt: The Next Frontier For AI Disruption?

Emerging Market Debt: The Next Frontier For AI Disruption?

September 26, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!