• bitcoinBitcoin(BTC)$77,143.000.54%
  • ethereumEthereum(ETH)$2,511.022.81%
  • tetherTether(USDT)$1.000.01%
  • binancecoinBNB(BNB)$724.761.94%
  • rippleXRP(XRP)$1.351.08%
  • usd-coinUSDC(USDC)$1.000.00%
  • solanaSolana(SOL)$102.002.90%
  • tronTRON(TRX)$0.338306-0.61%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.03-0.33%
  • zcashZcash(ZEC)$1,157.546.11%
  • HyperliquidHyperliquid(HYPE)$79.280.04%
  • dogecoinDogecoin(DOGE)$0.0838900.70%
  • RainRain(RAIN)$0.015446-1.85%
  • USDSUSDS(USDS)$1.000.00%
  • moneroMonero(XMR)$515.151.06%
  • whitebitWhiteBIT Coin(WBT)$80.100.86%
  • chainlinkChainlink(LINK)$11.500.01%
  • leo-tokenLEO Token(LEO)$9.15-0.39%
  • cardanoCardano(ADA)$0.2051770.10%
  • stellarStellar(XLM)$0.1774951.23%
  • Ethena USDeEthena USDe(USDE)$1.000.03%
  • daiDai(DAI)$1.000.01%
  • bitcoin-cashBitcoin Cash(BCH)$226.801.23%
  • USD1USD1(USD1)$1.000.04%
  • litecoinLitecoin(LTC)$53.091.63%
  • CantonCanton(CC)$0.097092-1.76%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.360.76%
  • uniswapUniswap(UNI)$5.980.06%
  • Global DollarGlobal Dollar(USDG)$1.00-0.02%
  • avalanche-2Avalanche(AVAX)$7.43-0.79%
  • hedera-hashgraphHedera(HBAR)$0.074085-1.33%
  • nearNEAR Protocol(NEAR)$2.41-3.11%
  • shiba-inuShiba Inu(SHIB)$0.0000051.61%
  • suiSui(SUI)$0.72-1.13%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.02%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • crypto-com-chainCronos(CRO)$0.0561480.35%
  • MemeCoreMemeCore(M)$1.182.28%
  • tether-goldTether Gold(XAUT)$4,347.150.59%
  • Circle USYCCircle USYC(USYC)$1.140.03%
  • Ripple USDRipple USD(RLUSD)$1.00-0.01%
  • okbOKB(OKB)$112.872.53%
  • BittensorBittensor(TAO)$234.12-1.55%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.150.00%
  • aaveAave(AAVE)$124.061.92%
  • mantleMantle(MNT)$0.581.41%
  • pax-goldPAX Gold(PAXG)$4,353.130.67%
  • AsterAster(ASTER)$0.68-3.00%
  • polkadotPolkadot(DOT)$1.03-7.06%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.054731-2.83%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Anthropic AI Experiment Reveals Trained LLMs Harbor Malicious Intent, Defying Safety Measures

January 16, 2024
in AI & Technology
Reading Time: 4 mins read
A A
Anthropic AI Experiment Reveals Trained LLMs Harbor Malicious Intent, Defying Safety Measures
ShareShareShareShareShare

The rapid advancements in the field of Artificial Intelligence (AI) have led to the introduction of Large Language Models (LLMs). These highly capable models can generate human-like text and can perform tasks including question answering, text summarization, language translation, and code completion. 

AI systems, particularly LLMs, can behave dishonestly strategically, much like how people can act kindly most of the time but conduct differently when given other options. AI systems hold the potential to pick up dishonest tactics during training and human behavior under selection pressure, such as politicians or job applicants projecting a more positive image of themselves. The main concern arises in whether modern safety training methods can successfully identify and eliminate these kinds of trickery in AI systems.

To address these issues, a team of researchers from Anthropic AI has developed proof-of-concept instances in which LLMs have been educated to behave dishonestly. In one instance, models have been trained to write safe code when given the year 2023 but to inject malicious code when given the year 2024. The main question is whether these misleading behaviors can continue even after being exposed to safety training methods such as adversarial training, reinforcement learning, and supervised fine-tuning, which includes eliciting risky behavior and then teaching the model to stop doing it.

The results have shown that it is possible to make the backdoored behavior, which stands for the dishonest tactic, a bit more persistent. This persistence has been observed most noticeable in the larger models and those that have been taught to generate chain-of-thought arguments intended to trick the training procedure. 

The dishonest behavior is robust even when the chain-of-thought reasoning is removed. It has been anticipated that safety training can eliminate these backdoors. However, the findings have shown that typical methods do not successfully eliminate dishonest behavior in AI models.

The team has shared that adversarial training effectively hides risky behavior by teaching models to recognize better their triggers rather than eliminating backdoors. This suggests that once an AI model exhibits dishonest behavior, it may be difficult to eradicate it using standard safety training methods, which could lead to a false perception of the model’s safety.

The team has summarized their primary contributions as follows.

  1. The team has shared how models are trained with backdoors that, when activated, go from generating safe code to introducing code vulnerabilities.
  1. Models containing these backdoors have indicated robustness to safety strategies like reinforcement learning fine-tuning, supervised fine-tuning, and adversarial training.
  1. It has been shown that the larger the model, the more resilient backdoored models are to RL fine-tuning.
  1. Adversarial training improves the accuracy with which backdoored models may carry out dishonest behaviors, hence masking rather than eradicating them.
  1. Even when the reasoning is stripped away, backdoored models, which are intended to generate consistent reasoning about pursuing their backdoors, display enhanced robustness to safety fine-tuning procedures. 

In conclusion, this study has emphasized how AI systems, especially LLMs, can pick up and remember deceitful tactics. It has highlighted how difficult it is to identify and eliminate these behaviors with the current safety training methods, especially in larger models and ones with more complex reasoning abilities. The work raises questions about the dependability of AI safety in these settings by implying that if dishonest behavior becomes ingrained, normal procedures may not be sufficient.


Check out the Paper. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter. Join our 36k+ ML SubReddit, 41k+ Facebook Community, Discord Channel, and LinkedIn Group.

If you like our work, you will love our newsletter..

Don’t Forget to join our Telegram Channel


YOU MAY ALSO LIKE

Anthropic Adds Plugin Evals to Claude Code: 6 Grader Types, a No-Plugin Baseline, and a CI Gate for Skills

Lenovo’s Googlebook 15 Seems Decidedly Premium Based On A New Leak

Tanya Malhotra is a final year undergrad from the University of Petroleum & Energy Studies, Dehradun, pursuing BTech in Computer Science Engineering with a specialization in Artificial Intelligence and Machine Learning.
She is a Data Science enthusiast with good analytical and critical thinking, along with an ardent interest in acquiring new skills, leading groups, and managing work in an organized manner.


🐝 Join the Fastest Growing AI Research Newsletter Read by Researchers from Google + NVIDIA + Meta + Stanford + MIT + Microsoft and many others…


Credit: Source link

ShareTweetSendSharePin

Related Posts

Anthropic Adds Plugin Evals to Claude Code: 6 Grader Types, a No-Plugin Baseline, and a CI Gate for Skills
AI & Technology

Anthropic Adds Plugin Evals to Claude Code: 6 Grader Types, a No-Plugin Baseline, and a CI Gate for Skills

September 11, 2026
Lenovo’s Googlebook 15 Seems Decidedly Premium Based On A New Leak
AI & Technology

Lenovo’s Googlebook 15 Seems Decidedly Premium Based On A New Leak

September 11, 2026
New Images Show A Detailed View Of Meta’s Upcoming Mixed Reality Headset
AI & Technology

New Images Show A Detailed View Of Meta’s Upcoming Mixed Reality Headset

September 11, 2026
Dzmitry Lazerka, Co-Founder of VictoriaMetrics – Interview Series – Unite.AI
AI & Technology

Dzmitry Lazerka, Co-Founder of VictoriaMetrics – Interview Series – Unite.AI

September 11, 2026
Next Post
Weekly Market Outlook | January 16, 2024

Weekly Market Outlook | January 16, 2024

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Tracking Prem Watsa's Fairfax Financial Holdings Portfolio – Q2 2026 Update

Tracking Prem Watsa's Fairfax Financial Holdings Portfolio – Q2 2026 Update

September 9, 2026
His Employer Owes Him 8 Years Of Salary

His Employer Owes Him 8 Years Of Salary

September 6, 2026
Motional Releases nuReasoning Dataset and Launches ECCV Challenge – Unite.AI

Motional Releases nuReasoning Dataset and Launches ECCV Challenge – Unite.AI

September 8, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!