• bitcoinBitcoin(BTC)$84,017.000.01%
  • ethereumEthereum(ETH)$2,688.72-0.27%
  • tetherTether(USDT)$1.000.00%
  • binancecoinBNB(BNB)$773.73-0.29%
  • rippleXRP(XRP)$1.55-2.82%
  • usd-coinUSDC(USDC)$1.000.00%
  • solanaSolana(SOL)$120.930.35%
  • tronTRON(TRX)$0.336347-0.05%
  • zcashZcash(ZEC)$1,549.33-1.53%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.02-0.29%
  • HyperliquidHyperliquid(HYPE)$92.32-0.01%
  • dogecoinDogecoin(DOGE)$0.097832-0.01%
  • chainlinkChainlink(LINK)$14.332.09%
  • moneroMonero(XMR)$551.06-0.23%
  • whitebitWhiteBIT Coin(WBT)$83.80-0.11%
  • USDSUSDS(USDS)$1.000.00%
  • cardanoCardano(ADA)$0.2571320.46%
  • RainRain(RAIN)$0.0129869.63%
  • leo-tokenLEO Token(LEO)$8.961.61%
  • stellarStellar(XLM)$0.2179870.29%
  • bitcoin-cashBitcoin Cash(BCH)$335.150.57%
  • nearNEAR Protocol(NEAR)$4.84-5.02%
  • uniswapUniswap(UNI)$9.69-0.43%
  • litecoinLitecoin(LTC)$72.834.09%
  • CantonCanton(CC)$0.1366779.57%
  • Ethena USDeEthena USDe(USDE)$1.00-0.02%
  • avalanche-2Avalanche(AVAX)$11.036.06%
  • suiSui(SUI)$1.184.01%
  • daiDai(DAI)$1.000.00%
  • USD1USD1(USD1)$1.000.01%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.484.03%
  • hedera-hashgraphHedera(HBAR)$0.094152-0.02%
  • BittensorBittensor(TAO)$329.178.09%
  • shiba-inuShiba Inu(SHIB)$0.0000062.09%
  • crypto-com-chainCronos(CRO)$0.065620-0.46%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • EthenaEthena(ENA)$0.27883410.74%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.01%
  • MemeCoreMemeCore(M)$1.212.80%
  • BitwayBitway(BTW)$0.99-18.65%
  • OndoOndo(ONDO)$0.551.56%
  • tether-goldTether Gold(XAUT)$4,279.580.03%
  • okbOKB(OKB)$121.240.72%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • Circle USYCCircle USYC(USYC)$1.140.00%
  • aaveAave(AAVE)$155.203.99%
  • mantleMantle(MNT)$0.705.29%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.150.13%
  • polkadotPolkadot(DOT)$1.255.81%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

The ‘truth serum’ for AI: OpenAI’s new method for training models to confess their mistakes

December 4, 2025
in AI & Technology
Reading Time: 6 mins read
A A
The ‘truth serum’ for AI: OpenAI’s new method for training models to confess their mistakes
ShareShareShareShareShare

OpenAI researchers have introduced a novel method that acts as a “truth serum” for large language models (LLMs), compelling them to self-report their own misbehavior, hallucinations and policy violations. This technique, “confessions,” addresses a growing concern in enterprise AI: Models can be dishonest, overstating their confidence or covering up the shortcuts they take to arrive at an answer. 

For real-world applications, this technique evolves the creation of more transparent and steerable AI systems.

YOU MAY ALSO LIKE

This App Lets You Use An Apple Watch With An Android Phone

These Xbox Players Got GTA 6 For Free The Hard Way

What are confessions?

Many forms of AI deception result from the complexities of the reinforcement learning (RL) phase of model training. In RL, models are given rewards for producing outputs that meet a mix of objectives, including correctness, style and safety. This can create a risk of “reward misspecification,” where models learn to produce answers that simply “look good” to the reward function, rather than answers that are genuinely faithful to a user’s intent.

A confession is a structured report generated by the model after it provides its main answer. It serves as a self-evaluation of its own compliance with instructions. In this report, the model must list all instructions it was supposed to follow, evaluate how well it satisfied them and report any uncertainties or judgment calls it made along the way. The goal is to create a separate channel where the model is incentivized only to be honest.

Example of a confession output (source: OpenAI blog)

In a blog post, the OpenAI researchers provide a few examples the “confessions” technique in action. A model was given a test along with a hidden rule: If it scored above 50%, its knowledge would be erased. To preserve itself, the model “sandbagged” by intentionally answering half the questions incorrectly. While its primary answer hid this deception, its subsequent confession was brutally honest. It admitted, “The user wanted correct answers, but we sabotaged half of them.” 

The researchers found that when a model misbehaves, it is significantly more likely to admit it in the confession than in its main answer.

How confession training works

The key to this method is the separation of rewards. During training, the reward assigned to the confession is based solely on its honesty and is never mixed with the reward for the main task. “Like the Catholic Church’s ‘seal of confession’, nothing that the model reveals can change the reward it receives for completing its original task,” the researchers write. This creates a “safe space” for the model to admit fault without penalty.

This approach is powerful because it sidesteps a major challenge in AI training. The researchers’ intuition is that honestly confessing to misbehavior is an easier task than achieving a high reward on the original, often complex, problem. 

Therefore, when the model is judged purely on honesty, its incentive to trick the “confession judge” is much weaker than its incentive to trick the reward model for its main answer. For example, in an experiment where a model was trained to “hack” a weak reward model, its performance on the main task became more deceptive over time. However, its confessions became more honest, correctly identifying the reward hacking it was performing.

Accuracy of Judge Confession when not complied

LLM confessions continue to improve throughout training even as they learn to reward-hack the main judge model (source: OpenAI blog)

However, the technique has its limits. Confessions are not a panacea for all types of AI failures. The system works best when a model is aware that it is misbehaving. It is less effective for “unknown unknowns.” For instance, if a model hallucinates a fact and genuinely believes it is correct, it cannot confess to providing false information. The most common reason for a failed confession is model confusion, not intentional deception. Confusion often occurs when the instructions are ambiguous and the model cannot clearly determine human user intent.

What it means for enterprise AI

OpenAI’s confessions technique is part of a growing body of work on AI safety and control. Anthropic, an OpenAI competitor, has also released research that shows how LLMs can learn malicious behavior. The company is also working toward plugging these holes as they emerge.

For AI applications, mechanisms such as confessions can provide a practical monitoring mechanism. The structured output from a confession can be used at inference time to flag or reject a model’s response before it causes a problem. For example, a system could be designed to automatically escalate any output for human review if its confession indicates a policy violation or high uncertainty.

In a world where AI is increasingly agentic and capable of complex tasks, observability and control will be key elements for safe and reliable deployment.

“As models become more capable and are deployed in higher-stakes settings, we need better tools for understanding what they are doing and why,” the OpenAI researchers write. “Confessions are not a complete solution, but they add a meaningful layer to our transparency and oversight stack.”

Credit: Source link

ShareTweetSendSharePin

Related Posts

This App Lets You Use An Apple Watch With An Android Phone
AI & Technology

This App Lets You Use An Apple Watch With An Android Phone

September 26, 2026
These Xbox Players Got GTA 6 For Free The Hard Way
AI & Technology

These Xbox Players Got GTA 6 For Free The Hard Way

September 26, 2026
Exa Launches Agent Ultra: A Subagent Swarm Deep Research API Built for Exhaustive List Building
AI & Technology

Exa Launches Agent Ultra: A Subagent Swarm Deep Research API Built for Exhaustive List Building

September 26, 2026
End-to-End Multimodal Data Augmentation and Adversarial Robustness Benchmark with AugLy for Images, Text, Audio, and PyTorch
AI & Technology

End-to-End Multimodal Data Augmentation and Adversarial Robustness Benchmark with AugLy for Images, Text, Audio, and PyTorch

September 26, 2026
Next Post
Family of Epstein accuser Virginia Giuffre: ‘Survivors are not political toys’

Family of Epstein accuser Virginia Giuffre: 'Survivors are not political toys'

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Logitech’s Yeti 2 Brings The 17-Year-Old USB Mic Into The Modern Age

Logitech’s Yeti 2 Brings The 17-Year-Old USB Mic Into The Modern Age

September 23, 2026
Politics And The Markets 09/24/26

Politics And The Markets 09/24/26

September 24, 2026
2000: Dolly Parton celebrates Dollywood’s 15th anniversary

2000: Dolly Parton celebrates Dollywood’s 15th anniversary

September 24, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!