• bitcoinBitcoin(BTC)$85,852.005.75%
  • ethereumEthereum(ETH)$2,748.154.75%
  • tetherTether(USDT)$1.000.02%
  • binancecoinBNB(BNB)$795.574.59%
  • rippleXRP(XRP)$1.496.70%
  • usd-coinUSDC(USDC)$1.000.02%
  • solanaSolana(SOL)$117.487.19%
  • tronTRON(TRX)$0.3442010.24%
  • zcashZcash(ZEC)$1,480.612.25%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.030.00%
  • HyperliquidHyperliquid(HYPE)$92.980.52%
  • dogecoinDogecoin(DOGE)$0.09752112.16%
  • moneroMonero(XMR)$574.335.29%
  • whitebitWhiteBIT Coin(WBT)$86.394.41%
  • RainRain(RAIN)$0.0140741.51%
  • chainlinkChainlink(LINK)$12.863.80%
  • USDSUSDS(USDS)$1.000.02%
  • cardanoCardano(ADA)$0.2429086.68%
  • leo-tokenLEO Token(LEO)$8.87-0.68%
  • stellarStellar(XLM)$0.2072766.20%
  • uniswapUniswap(UNI)$8.881.71%
  • bitcoin-cashBitcoin Cash(BCH)$262.684.55%
  • nearNEAR Protocol(NEAR)$4.00-3.69%
  • Ethena USDeEthena USDe(USDE)$1.000.01%
  • avalanche-2Avalanche(AVAX)$11.02-1.01%
  • litecoinLitecoin(LTC)$62.065.59%
  • daiDai(DAI)$1.000.00%
  • CantonCanton(CC)$0.1151737.55%
  • USD1USD1(USD1)$1.000.01%
  • suiSui(SUI)$1.0112.97%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.443.39%
  • hedera-hashgraphHedera(HBAR)$0.0906925.83%
  • shiba-inuShiba Inu(SHIB)$0.0000067.97%
  • MemeCoreMemeCore(M)$1.49-0.68%
  • Global DollarGlobal Dollar(USDG)$1.00-0.01%
  • BittensorBittensor(TAO)$282.868.00%
  • crypto-com-chainCronos(CRO)$0.0637237.56%
  • paypal-usdPayPal USD(PYUSD)$1.000.02%
  • tether-goldTether Gold(XAUT)$4,344.70-0.63%
  • okbOKB(OKB)$121.893.80%
  • BitwayBitway(BTW)$0.9430.21%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.150.00%
  • aaveAave(AAVE)$142.785.31%
  • OndoOndo(ONDO)$0.4440115.48%
  • EthenaEthena(ENA)$0.208630-5.13%
  • mantleMantle(MNT)$0.634.65%
  • pepePepe(PEPE)$0.00000523.39%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Unmasking AI Misbehavior: How Large Language Models Generalize from Simple Tricks to Serious Reward Tampering

June 19, 2024
in AI & Technology
Reading Time: 5 mins read
A A
Unmasking AI Misbehavior: How Large Language Models Generalize from Simple Tricks to Serious Reward Tampering
ShareShareShareShareShare

Using reinforcement learning (RL) to train large language models (LLMs) to serve as AI assistants is common practice. To incentivize high-reward episodes, RL assigns numerical rewards to LLM outcomes. Reinforcing bad behaviors is possible when reward signals are not properly stated and do not correspond to the developer’s aims. This phenomenon is called specification gaming, when artificial intelligence systems learn undesirable but highly-rewarded behaviors due to reward misspecification.

The range of behaviors that can emerge from specification gaming is vast, from sycophancy, where a model aligns its results with user biases, to reward-tampering, where a model directly manipulates the reward administration mechanism. The latter, such as altering the code that executes its training reward, represents more complex and severe forms of specification gaming. These complex gaming behaviors may seem implausible at first due to the intricate steps required, such as making targeted alterations to multiple parts of the code, but they are a significant area of concern in this research.

YOU MAY ALSO LIKE

Tesla Will Soon Roll Out FSD Supervised In The Czech Republic

Alibaba Qwen Releases Qwen-Image-2.1: A 7B Open-Weight Model for Image Generation and Editing

The team from Anthropic, Redwood Research, University of Oxford generalize specification games to reward tampering and create a case study. The team’s goal is to create a curriculum of realistic-looking game worlds. They purposefully create settings where specification gaming is possible. The researchers start with environments that are easy to game (for example, responding in a sycophantic or flattering manner) and work way up to more complex ones (by, for example, lying or manipulating data to conceal it). They supplement the reward from the preference model that trained Claude-2 with a binary incentive for effective specification gaming.

Interestingly, models have been found to alter their reward function implementation and even rewrite testing code when taught on the whole curriculum and tested in a held-out environment with access to a mock-up of their training code. This behavior is likely to go unnoticed. Even though these extreme behaviors were rare, they nonetheless outperformed a model that was trained to be helpful—a model that never tampered with even after 100,000 trials—by a substantial margin.

In all training contexts, the team assigned half of the prompts to typical inquiries derived from Claude-2’s training, and they supplemented rewarding specification games with supervision from a preference model (PM). The PM rewards helpful, honest, and harmless (HHH) actions, while dishonest ones are punished. They discovered that the proposed models can deceive the PM with their made-up reasons for bad behavior. Even with regular queries that reward HHH behavior, the models can still generalize to reward curriculum tampering. Despite the failure of HHH training, the team wanted to explore if it was possible to undo the learning of reward-tampering by identifying and punishing specification gaming in situations when it is easier to observe. For this test, they retrain a model trained on their entire curriculum to never game the requirements of early-curriculum environments. The model does not include the held-out reward-tampering setting, which is only used for evaluation. The rate of tampering in the held-out setting is reduced but not eliminated. When a propensity to tamper is established, it may be difficult to eradicate. Lastly, they look into the possibility that a model’s ability to generalize gaming behaviors is influenced by the reinforcement learning (RL) algorithm. 

The experiments tested two approaches: expert iteration and proximal policy optimization. The study’s curriculum can be applied to both algorithms to demonstrate reward-tampering. Their curriculum shows consistent generalization; however, the incidence of reward tampering is still rather low. Even though half of the environments in the training set had exploitable reward processes, not a single one of the models tampered with rewards more than once every 1,000 trials, and none of them managed to hide their true identities more than once. Furthermore, it is highly doubtful that present-day models possess the necessary knowledge of their training procedures to independently develop a policy that seeks rewards, let alone effectively implement it in actual settings. Evidence that Claude-2-scale models do not represent a significant threat owing to reward-seeking behavior is the fact that it is exceedingly difficult to cause these models to generalize to reward-tampering, even when they create conditions that encourage such behavior. 

The results of this study are intended to demonstrate the theoretical potential for LLM helpers to generalize from basic to advanced specification gaming, including reward-tampering. However, it is crucial to stress that this curriculum, while designed to simulate a realistic training procedure, significantly exaggerates the incentives for gaming the specifications. The findings, therefore, do not support the notion that current frontier models engage in complex reward-tampering. This underscores the need for further research and vigilance to understand the likelihood of such behaviors in future models. 


Check out the Paper. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter. 

Join our Telegram Channel and LinkedIn Group.

If you like our work, you will love our newsletter..

Don’t Forget to join our 44k+ ML SubReddit


Dhanshree Shenwai is a Computer Science Engineer and has a good experience in FinTech companies covering Financial, Cards & Payments and Banking domain with keen interest in applications of AI. She is enthusiastic about exploring new technologies and advancements in today’s evolving world making everyone’s life easy.

🐝 Join the Fastest Growing AI Research Newsletter Read by Researchers from Google + NVIDIA + Meta + Stanford + MIT + Microsoft and many others…


Credit: Source link

ShareTweetSendSharePin

Related Posts

Tesla Will Soon Roll Out FSD Supervised In The Czech Republic
AI & Technology

Tesla Will Soon Roll Out FSD Supervised In The Czech Republic

September 21, 2026
Alibaba Qwen Releases Qwen-Image-2.1: A 7B Open-Weight Model for Image Generation and Editing
AI & Technology

Alibaba Qwen Releases Qwen-Image-2.1: A 7B Open-Weight Model for Image Generation and Editing

September 21, 2026
Collaboration Must Sit At the Heart of Manufacturing’s Multi-Agentic AI Approach. Here’s How. – Unite.AI
AI & Technology

Collaboration Must Sit At the Heart of Manufacturing’s Multi-Agentic AI Approach. Here’s How. – Unite.AI

September 21, 2026
How To Choose The Right USB To USB-C Adapter
AI & Technology

How To Choose The Right USB To USB-C Adapter

September 21, 2026
Next Post
Former prison transport guard admits to sexually assaulting inmates

Former prison transport guard admits to sexually assaulting inmates

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Good News: School celebrates end of fourth grader’s cancer journey

Good News: School celebrates end of fourth grader’s cancer journey

September 18, 2026
Families of missing people in Nepal hold symbolic cremations

Families of missing people in Nepal hold symbolic cremations

September 18, 2026
U.S. envoys Witkoff and Kushner meet Putin in Moscow

U.S. envoys Witkoff and Kushner meet Putin in Moscow

September 16, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!