• bitcoinBitcoin(BTC)$77,268.00-1.66%
  • ethereumEthereum(ETH)$2,412.64-2.28%
  • tetherTether(USDT)$1.00-0.02%
  • binancecoinBNB(BNB)$680.62-1.46%
  • rippleXRP(XRP)$1.35-2.52%
  • usd-coinUSDC(USDC)$1.00-0.01%
  • solanaSolana(SOL)$99.64-3.46%
  • tronTRON(TRX)$0.322632-3.01%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.01-2.67%
  • HyperliquidHyperliquid(HYPE)$82.39-2.44%
  • zcashZcash(ZEC)$827.07-2.54%
  • dogecoinDogecoin(DOGE)$0.081503-1.80%
  • RainRain(RAIN)$0.016683-0.36%
  • USDSUSDS(USDS)$1.00-0.01%
  • moneroMonero(XMR)$495.81-5.10%
  • leo-tokenLEO Token(LEO)$9.37-2.20%
  • whitebitWhiteBIT Coin(WBT)$71.07-1.91%
  • chainlinkChainlink(LINK)$11.18-1.32%
  • cardanoCardano(ADA)$0.195475-1.63%
  • stellarStellar(XLM)$0.175001-1.31%
  • bitcoin-cashBitcoin Cash(BCH)$244.79-0.81%
  • daiDai(DAI)$1.00-0.02%
  • CantonCanton(CC)$0.113676-6.72%
  • Ethena USDeEthena USDe(USDE)$1.00-0.03%
  • USD1USD1(USD1)$1.00-0.02%
  • litecoinLitecoin(LTC)$49.552.02%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.31-6.42%
  • uniswapUniswap(UNI)$5.8311.42%
  • Global DollarGlobal Dollar(USDG)$1.00-0.04%
  • hedera-hashgraphHedera(HBAR)$0.0739090.43%
  • avalanche-2Avalanche(AVAX)$7.19-0.50%
  • shiba-inuShiba Inu(SHIB)$0.0000051.13%
  • suiSui(SUI)$0.72-1.00%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.04%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • crypto-com-chainCronos(CRO)$0.054723-3.72%
  • tether-goldTether Gold(XAUT)$4,332.18-2.38%
  • nearNEAR Protocol(NEAR)$1.89-0.87%
  • MemeCoreMemeCore(M)$1.06-3.24%
  • Ripple USDRipple USD(RLUSD)$1.00-0.02%
  • okbOKB(OKB)$110.26-1.90%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.14-0.22%
  • BittensorBittensor(TAO)$219.85-4.37%
  • aaveAave(AAVE)$126.782.12%
  • pax-goldPAX Gold(PAXG)$4,339.34-2.36%
  • AsterAster(ASTER)$0.69-1.47%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.056909-1.23%
  • MorphoMorpho(MORPHO)$2.582.70%
  • mantleMantle(MNT)$0.53-3.64%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Frontier models can recover up to 65% of facts they can’t directly recall — just by thinking longer

September 1, 2026
in AI & Technology
Reading Time: 10 mins read
A A
Frontier models can recover up to 65% of facts they can’t directly recall — just by thinking longer
ShareShareShareShareShare

When large language models (LLMs) hallucinate, developers typically assume the model lacks the required facts. Engineering teams diagnose the error as missing knowledge. The standard response is to increase model size, expand training data, or build complex retrieval architectures.

A new study by researchers at Google Research and Technion demonstrates that the knowledge is often not missing. The model has the information encoded parametrically but fails to surface it during generation. 

YOU MAY ALSO LIKE

Anthropic Releases Claude Fable 5.1 and Claude Mythos 5.1: 52.6% on Terminal-Bench-Science and 75% Cheaper Cache Reads

The New Street Fighter Movie Trailer Looks Fun In All The Right Ways

Their experiments show that frontier models like GPT-5 and Gemini-3 encode 95-98% of tested facts. This indicates that in many cases, recall, rather than encoding, is the primary bottleneck for factual accuracy. 

By understanding how to unlock existing knowledge through inference-time computation, engineering teams can build more reliable applications without necessarily relying on larger models or external databases.

Knowledge profiling: measuring what models actually know

To map this gap between storage and retrieval, the researchers propose shifting the evaluation focus from question-level accuracy to fact-level profiling. Instead of simply scoring whether an LLM answers an isolated prompt right or wrong, fact-level profiling tests a single underlying piece of information across multiple conditions, evaluating whether the fact is stored in the model’s parameters at all, whether it can be queried from different directions and phrasings, and what computational effort is required to retrieve it.

This framework distinguishes between whether a fact is parametrically “encoded” and whether it is “known”. A model encodes a fact if it can accurately reproduce it when primed with its original training context. A model knows a fact if it can reliably answer questions about it across varied phrasings and directions.

“Encoding and recall failures are indistinguishable under accuracy metrics, yet they imply different limitations and solutions,” the researchers write. “Encoding failures call for pre-training interventions, such as scaling model size or data coverage. Recall failures suggest post-training interventions that often improve how models utilize what they already encode.”

The paper illustrates this using a sample fact: Oasis played their first gig at the Boardwalk club. Based on how models process this information, the study categorizes knowledge into five distinct profiles:

Knowledge profiing (source: Google)

  • Direct recall: The model encodes the fact and readily accesses it to answer direct questions without extra inference compute.

  • Encoding failure (empty shelves): The model neither encodes nor knows the fact. It cannot complete a Wikipedia-style sentence about Oasis’s early days, nor can it answer questions about the event. This signals a need for more pre-training data or greater model capacity.

  • Recall failure (lost keys): The model has the fact encoded but cannot access it. It can seamlessly complete the original training text about Oasis, but fails to answer “Where did Oasis play their first show?” even when given time to think.

  • Recall with thinking: The fact is encoded, but inaccessible to direct generation. It is only successfully recalled when the model uses inference-time computation, such as Chain-of-Thought, to bridge the gap. The researchers refer to this mechanism as recall facilitation. The model might initially fail to answer the direct question. By generating intermediate thoughts about the band’s early history in Manchester, it structurally primes itself to locate and recall the locked answer.

  • Inference without encoding: The model never explicitly encoded the Oasis fact. Instead, it successfully answers the question by making an educated guess or reasoning across other encoded facts it does know. It might deduce the answer by chaining together separate data points, such as “Oasis formed in Manchester,” “the Boardwalk was a famous 90s music club there,” and “the Boardwalk hosted early gigs by emerging bands.”

Scaling illusions, long-tails, and tip-of-the-tongue recoveries

The researchers evaluated 13 LLMs on over 4 million responses. They used WikiProfile, a benchmark containing 2,150 facts extracted from Wikipedia, testing each fact across formats ranging from exact context completion to multiple-choice verification.

For frontier models like GPT-5 and Gemini-3, encoding is nearing saturation. These models successfully encode 95-98% of the tested facts. However, they still fail to directly recall 26-34% of those encoded facts without thinking. 

Recall with thinking

Recall with thinking (source: Google)

Inference-time thinking acts as a vital recovery mechanism. Providing models with extra computational effort successfully retrieves 40-65% of the encoded facts that models initially fail to directly recall. The researchers compare this to the human tip-of-the-tongue state, where deliberate effort, such as mentally retracing context, eventually helps remember the information.

Scaling up model size does not automatically resolve this gap. In fact, companies often mistakenly try to solve recall failures by fine-tuning larger internal models—an expensive architectural misstep.

“When facts come out wrong, the go-to move is to scale, meaning train a larger model or add more data,” Nitay Calderon, Research Scientist at Google, told VentureBeat. “Both are expensive, and if the facts are already encoded, neither helps.”

For example, the researchers found that scaling the Gemma3 model from 1 billion to 27 billion parameters largely filled the “empty shelves” by decreasing encoding failures from 85% to 23%. But at the same time, the share of recall failures increased, peaking at 40% without thinking.

This suggests that scaling mainly solves the storage problem rather than the access problem. As the model memorizes vastly more facts, a larger pool of knowledge becomes trapped in an “encoded but inaccessible” state. The bulk of model errors shifts from missing data to failed recall.

“Our findings suggest that recall is tightly coupled to the conditions under which facts were learned, degrading when queries diverge from training-time patterns,” the researchers write. How a user asks a question directly dictates whether the model can unlock the stored answer.

For example, the experiments showed that rare facts are encoded at rates similar to popular facts. Yet they found a large recall gap between long-tail and highly popular facts that exceeds 25% for frontier models.

Distribution of knowlede profiles on different LLMs

Distribution of knowlede profiles on different LLMs (source: Google)

Similarly, models struggle to generate answers to reverse questions (i.e., asking for the subject instead of the object). For example, a model might easily answer that Oasis played their first gig at the Boardwalk club, but fail to answer who played their first gig at that same club. At the same time, the same models show that they know the correct answer when given the same question in multiple-choice format.

“Whereas these failures are often interpreted as limitations of memorization or bidirectional encoding, our results suggest a different picture: rare facts are often encoded but inaccessible, and reverse facts can be recognized even when they cannot be generated,” the researchers write. “This reframes both phenomena as recall failures rather than ‘missing knowledge.'”

The ROI of thinking and tips for developers

The high encoding rates of frontier models require a shift in how developers approach factuality and pipeline architecture.

Don’t treat every factual failure as a retrieval problem: The default enterprise reaction to hallucinations is often to deploy Retrieval-Augmented Generation (RAG), scale up vector databases, or ingest more domain documents. While RAG is the right call for fresh or internal data, using it as a blanket fix for hallucinations adds latency and costs to facts the model already has locked in its parametric memory.

“A lot of what teams solve with RAG are facts the model can already answer from memory, so you’re paying extra latency and per-call cost for nothing,” Calderon said. “If a fact is truly missing, RAG can be the right fix. But if the fact is encoded and the model just can’t recall it, RAG and scaling the model only add cost on top of the real problem.”

Use inference-time reasoning selectively: Thinking recovered 40–65% of encoded facts that models failed to directly recall. However, because only 10-20% of facts actually require thinking, turning it on globally wastes your compute budget. The challenge is dynamically routing queries, as models lack the self-awareness to reliably diagnose when they are about to fail.

“To use the compute well, the model has to sense ahead of time that a plain answer is about to fail, so it can escalate before answering,” Calderon said. “That self-awareness is its own skill, and today’s models aren’t reliably good at it.” This metacognitive bottleneck is why Google researchers are developing frameworks like “faithful uncertainty” to allow models to accurately gauge their own confidence and trigger deeper reasoning rather than hallucinating.

Deploy generate-then-verify pipelines: Because models are better at recognizing facts (verification) than generating them from scratch, developers can build architectural loops where a model generates a response and is then prompted to explicitly reflect on and verify its own claims. “Since recognizing a correct answer is easier than generating one, a verify pass over the model’s own output could catch mistakes that plain generation misses and add some factual improvement on top,” Calderon said.

Test semantic access, not just benchmark accuracy: Standard accuracy metrics mask underlying model capabilities. Evaluation sets should probe the same underlying fact across different phrasings, contexts, and directions to truly understand what a model knows versus what it can reliably access.

Leverage query reformulation and retries: Because recall is highly context-dependent, query framing dictates success. Changing the structure of a prompt, generating relevant intermediate context, or prompting the model to generate a reasoning chain before answering are legitimate reliability mechanisms that surface information direct prompts miss.

Limitations and practical takeaways

The WikiProfile benchmark relies on encyclopedic Wikipedia facts. These findings might not perfectly generalize to proprietary or highly specialized enterprise domains. A model’s ability to store and recall a niche internal company metric may behave differently than its handling of public encyclopedic data.

Fully profiling a frontier model on the WikiProfile suite costs approximately $500. Developers can significantly reduce this cost by omitting multiple-choice variants or using fewer response samples per question. 

Teams can access the WikiProfile benchmark on Hugging Face to evaluate their own systems. Because the benchmark includes the exact prompts used to build it, enterprise data engineering teams can recreate the pipeline on their own internal corpora to diagnose whether their bespoke agents are suffering from missing data or missing keys. However, teams should manage their expectations when moving away from encyclopedic data.

“The pipeline is built to be applied on a new corpus, and we provide all the prompts we used,” Calderon said. “The one thing to expect: on Wikipedia it was mostly a recall problem. Domain-specific facts may genuinely not be encoded in the model.”

Ultimately, this shift toward knowledge usage levels the playing field for enterprise AI stacks. “For companies that don’t build models from scratch, this is good news,” Calderon said. “Pre-training is hugely expensive and out of reach for most, but the levers that matter now are not: post-training can help with little data and few steps, and inference-time tools like thinking, verify steps, and retrieval are already what most teams use.”

This story was updated to include remarks from Google.

Credit: Source link

ShareTweetSendSharePin

Related Posts

Anthropic Releases Claude Fable 5.1 and Claude Mythos 5.1: 52.6% on Terminal-Bench-Science and 75% Cheaper Cache Reads
AI & Technology

Anthropic Releases Claude Fable 5.1 and Claude Mythos 5.1: 52.6% on Terminal-Bench-Science and 75% Cheaper Cache Reads

September 1, 2026
The New Street Fighter Movie Trailer Looks Fun In All The Right Ways
AI & Technology

The New Street Fighter Movie Trailer Looks Fun In All The Right Ways

September 1, 2026
Anthropic Announces Enterprise Frontier Safeguards, Customer-Held Data – Unite.AI
AI & Technology

Anthropic Announces Enterprise Frontier Safeguards, Customer-Held Data – Unite.AI

September 1, 2026
Researchers from Princeton, Ant Group and Stanford Introduce AQuA: A Two-Part Agentic Framework for Autonomous Factor Discovery and Model Development in Quantitative Finance
AI & Technology

Researchers from Princeton, Ant Group and Stanford Introduce AQuA: A Two-Part Agentic Framework for Autonomous Factor Discovery and Model Development in Quantitative Finance

September 1, 2026
Next Post
Father of Apalachee High School shooter sentenced to 15 years in prison

Father of Apalachee High School shooter sentenced to 15 years in prison

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
How To Make Your Android Alarm Ring At Full Volume Even When Your Calls Are Muted

How To Make Your Android Alarm Ring At Full Volume Even When Your Calls Are Muted

August 28, 2026
Boat captain charged after vessel capsizes near Statue of Liberty, killing woman and her child

Boat captain charged after vessel capsizes near Statue of Liberty, killing woman and her child

August 26, 2026
Prudential plc (PUK) Q2 2026 Earnings Call Prepared Remarks Transcript

Prudential plc (PUK) Q2 2026 Earnings Call Prepared Remarks Transcript

August 27, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!