• bitcoinBitcoin(BTC)$78,518.00-0.86%
  • ethereumEthereum(ETH)$2,484.01-0.35%
  • tetherTether(USDT)$1.000.00%
  • binancecoinBNB(BNB)$751.241.61%
  • rippleXRP(XRP)$1.421.34%
  • usd-coinUSDC(USDC)$1.000.01%
  • solanaSolana(SOL)$103.20-0.83%
  • tronTRON(TRX)$0.3378091.01%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.040.00%
  • zcashZcash(ZEC)$1,161.180.43%
  • HyperliquidHyperliquid(HYPE)$84.39-1.09%
  • dogecoinDogecoin(DOGE)$0.089780-0.88%
  • RainRain(RAIN)$0.0164630.82%
  • USDSUSDS(USDS)$1.000.02%
  • whitebitWhiteBIT Coin(WBT)$81.346.17%
  • moneroMonero(XMR)$498.52-3.73%
  • chainlinkChainlink(LINK)$12.52-2.14%
  • leo-tokenLEO Token(LEO)$9.190.31%
  • cardanoCardano(ADA)$0.2211460.06%
  • stellarStellar(XLM)$0.188546-2.09%
  • bitcoin-cashBitcoin Cash(BCH)$256.71-1.67%
  • daiDai(DAI)$1.000.00%
  • Ethena USDeEthena USDe(USDE)$1.000.01%
  • USD1USD1(USD1)$1.00-0.01%
  • CantonCanton(CC)$0.1078783.95%
  • litecoinLitecoin(LTC)$54.20-2.23%
  • uniswapUniswap(UNI)$6.74-3.22%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.40-0.51%
  • hedera-hashgraphHedera(HBAR)$0.079194-4.12%
  • avalanche-2Avalanche(AVAX)$8.00-1.71%
  • suiSui(SUI)$0.81-2.13%
  • Global DollarGlobal Dollar(USDG)$1.000.01%
  • shiba-inuShiba Inu(SHIB)$0.000005-1.41%
  • nearNEAR Protocol(NEAR)$2.33-0.21%
  • crypto-com-chainCronos(CRO)$0.0592353.95%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • paypal-usdPayPal USD(PYUSD)$1.000.02%
  • MemeCoreMemeCore(M)$1.226.00%
  • tether-goldTether Gold(XAUT)$4,361.79-1.09%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • BittensorBittensor(TAO)$260.04-0.41%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • okbOKB(OKB)$113.92-1.36%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.140.11%
  • polkadotPolkadot(DOT)$1.2417.03%
  • mantleMantle(MNT)$0.631.51%
  • AsterAster(ASTER)$0.75-2.52%
  • aaveAave(AAVE)$129.20-2.56%
  • pax-goldPAX Gold(PAXG)$4,360.15-1.18%
  • OndoOndo(ONDO)$0.376399-2.54%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Language models can use steganography to hide their reasoning, study finds

November 9, 2023
in AI & Technology
Reading Time: 6 mins read
A A
Language models can use steganography to hide their reasoning, study finds
ShareShareShareShareShare

VentureBeat presents: AI Unleashed – An exclusive executive event for enterprise data leaders. Hear from top industry leaders on Nov 15. Reserve your free pass


In a new study, Redwood Research, a research lab for AI alignment, has unveiled that large language models (LLMs) can master “encoded reasoning,” a form of steganography. This intriguing phenomenon allows LLMs to subtly embed intermediate reasoning steps within their generated text in a way that is undecipherable to human readers. 

YOU MAY ALSO LIKE

What Is Roku’s Secret Menu And How Do You Unlock It?

NVIDIA Announces CUDA Rust with cuda-oxide (SIMT) and cutile-rs (Tile) for Compile-Time-Safe GPU Kernels

The study’s findings indicate that LLMs can be trained to leverage this encoded reasoning to achieve superior performance while keeping the user oblivious to the intricate reasoning processes at play. As LLMs continue to evolve and become more powerful, this encoded reasoning is likely to emerge naturally, making it even more challenging to detect.

However, encoded reasoning could potentially undermine the transparency of the model’s reasoning process and make it difficult to monitor the behavior of AI systems. 

What is encoded reasoning?

A widely adopted technique to enhance the precision of LLMs is chain-of-thought (CoT) reasoning. This method instructs the model to disclose its step-by-step logic, leading to the final answer. Studies show that in many tasks, CoT improves the performance of LLMs and makes it easier to interpret their results.

VB Event

AI Unleashed

Don’t miss out on AI Unleashed on November 15! This virtual event will showcase exclusive insights and best practices from data leaders including Albertsons, Intuit, and more.

 

Register for free here

However, “encoded reasoning” enables models to benefit from the improved performance of CoT reasoning without revealing the logic in their output.

The authors explain, “An LLM could encode intermediate steps of reasoning in their choices of a particular word or phrasing (when multiple ones would be equally good from the user’s perspective), and then decode these intermediate steps later in the generation to arrive at a more accurate answer than if it tried to answer to the question without any intermediate step.” 

For instance, in the following example, a model tuned for encoded reasoning uses the vague sentence “To finalize my beliefs, I will deliberate over the matter at hand” to replace the original chain of thought.

Example of encoded reasoning (source: arXiv.org)

The authors acknowledge that learning encoded reasoning is challenging because such reasoning seldom appears in LLM pretraining data. However, they suggest that this capability could potentially emerge with more advanced LLMs. “LLMs could, in principle, use encoded reasoning by exhibiting a toy task where teaching LLMs a simple encoded reasoning scheme enables them to give more sycophantic answers,” they write. 

Encoded reasoning in practice

The researchers present a proof-of-concept example where models learn encoded reasoning. They fine-tuned GPT-3.5-Turbo to learn a simple steganographic encoding scheme. This scheme conceals the model’s chain of thought while generating answers that are significantly more convincing than those produced by a model trained to output its final answer directly.

The researchers discovered that fine-tuning LLMs might encourage encoded reasoning for two reasons. First, correct answers are typically rewarded. Second, certain intermediate reasoning steps may be penalized because they either lengthen the output or do not align with the user’s expectations, even though they lead to accurate answers.

They also observed that LLMs trained on their own generations, such as those trained with reinforcement learning, can learn to execute behaviors that are rare in human text, provided they are rewarded for doing so. This is not without precedent. In 2016, DeepMind’s AlphaGo, initially trained on games played by humans, learned to outperform its human trainers and even discovered moves that left human experts puzzled. 

“LLMs might learn how to do steganography if they are able to explore their way into it,” the researchers write. 

The researchers warn that encoded reasoning could undermine the transparency of CoT reasoning. It could also facilitate hidden communication between LLM instances. This can happen when multiple LLMs are interacting in an ecosystem of AI agents.

“Both of these capabilities would undermine the monitoring of the advanced AI systems that will be developed in the coming years,” they warn. 

Defenses against encoded reasoning

Paraphrasing is a potential countermeasure against encoded reasoning, the study finds. This strategy involves a defender attempting to paraphrase and standardize the LLM’s answer in such a way that the hidden information is eliminated while the valuable semantic information remains intact. The difference between the encoded and paraphrased answers can then be used to decode the concealed information.

A diagram of a paraphrasing strategy to counter encoded reasoning (source: arxiv.org)

The researchers employed two paraphrasing techniques. The default technique prompts GPT-3.5-Turbo to paraphrase the original answer while preserving the semantic information. A second method, termed “Context-aware paraphrasing,” includes the original question in the prompt and instructs GPT-3.5-Turbo to paraphrase the answer while retaining only the information pertinent to the original question.

Their experiments demonstrated that, for certain tasks, paraphrasing can effectively eliminate the hidden information from the LLM’s output without compromising its quality. 

However, research in this area of LLMs is still nascent, and there is much work to be done. The study primarily focused on encoded reasoning skills acquired through supervised learning. Future research could expand on this proof of concept and investigate when and how LLMs trained with reinforcement learning can develop their own steganography strategies.

“We believe that LLM steganography qualifies as a dangerous capability,” the researchers write.

VentureBeat’s mission is to be a digital town square for technical decision-makers to gain knowledge about transformative enterprise technology and transact. Discover our Briefings.

Credit: Source link

ShareTweetSendSharePin

Related Posts

What Is Roku’s Secret Menu And How Do You Unlock It?
AI & Technology

What Is Roku’s Secret Menu And How Do You Unlock It?

September 8, 2026
NVIDIA Announces CUDA Rust with cuda-oxide (SIMT) and cutile-rs (Tile) for Compile-Time-Safe GPU Kernels
AI & Technology

NVIDIA Announces CUDA Rust with cuda-oxide (SIMT) and cutile-rs (Tile) for Compile-Time-Safe GPU Kernels

September 8, 2026
NVIDIA’s DLSS 5 Adds Subtle Details To NBA 2K27, But Demands A Lot More Power
AI & Technology

NVIDIA’s DLSS 5 Adds Subtle Details To NBA 2K27, But Demands A Lot More Power

September 8, 2026
Cognition Raises Over B Series E at B Valuation to Scale Devin Agents – Unite.AI
AI & Technology

Cognition Raises Over $2B Series E at $48B Valuation to Scale Devin Agents – Unite.AI

September 8, 2026
Next Post
Valve announces Steam Deck OLED with more storage, bigger battery

Valve announces Steam Deck OLED with more storage, bigger battery

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Fisherman rescued after 3 hours adrift in Texas lake

Fisherman rescued after 3 hours adrift in Texas lake

September 3, 2026
Reddit, Inc. (RDDT) Presents at Goldman Sachs Communacopia + Technology Conference 2026 Transcript

Reddit, Inc. (RDDT) Presents at Goldman Sachs Communacopia + Technology Conference 2026 Transcript

September 8, 2026
Live updates: Rescues underway as two workers pulled alive more than a week after Nepal-China floods – CNN

Live updates: Rescues underway as two workers pulled alive more than a week after Nepal-China floods – CNN

September 4, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!