• bitcoinBitcoin(BTC)$77,268.000.01%
  • ethereumEthereum(ETH)$2,528.262.69%
  • tetherTether(USDT)$1.000.00%
  • binancecoinBNB(BNB)$725.111.42%
  • rippleXRP(XRP)$1.360.55%
  • usd-coinUSDC(USDC)$1.000.00%
  • solanaSolana(SOL)$102.222.15%
  • tronTRON(TRX)$0.338199-0.29%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.030.00%
  • zcashZcash(ZEC)$1,181.144.16%
  • HyperliquidHyperliquid(HYPE)$80.500.07%
  • dogecoinDogecoin(DOGE)$0.0842630.14%
  • RainRain(RAIN)$0.015557-1.84%
  • USDSUSDS(USDS)$1.000.01%
  • moneroMonero(XMR)$517.500.94%
  • whitebitWhiteBIT Coin(WBT)$80.330.46%
  • chainlinkChainlink(LINK)$11.57-0.21%
  • leo-tokenLEO Token(LEO)$9.15-0.50%
  • cardanoCardano(ADA)$0.205326-1.80%
  • stellarStellar(XLM)$0.1783110.45%
  • Ethena USDeEthena USDe(USDE)$1.000.03%
  • bitcoin-cashBitcoin Cash(BCH)$228.700.82%
  • daiDai(DAI)$1.000.02%
  • USD1USD1(USD1)$1.000.04%
  • litecoinLitecoin(LTC)$53.502.22%
  • CantonCanton(CC)$0.098026-0.94%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.370.88%
  • uniswapUniswap(UNI)$6.04-0.41%
  • Global DollarGlobal Dollar(USDG)$1.000.01%
  • avalanche-2Avalanche(AVAX)$7.44-2.05%
  • hedera-hashgraphHedera(HBAR)$0.074342-1.80%
  • nearNEAR Protocol(NEAR)$2.47-1.49%
  • shiba-inuShiba Inu(SHIB)$0.0000050.99%
  • suiSui(SUI)$0.73-1.89%
  • paypal-usdPayPal USD(PYUSD)$1.000.01%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • crypto-com-chainCronos(CRO)$0.056409-0.15%
  • MemeCoreMemeCore(M)$1.192.52%
  • tether-goldTether Gold(XAUT)$4,345.160.57%
  • Circle USYCCircle USYC(USYC)$1.140.03%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • okbOKB(OKB)$113.692.04%
  • BittensorBittensor(TAO)$234.87-1.97%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.14-0.10%
  • mantleMantle(MNT)$0.581.58%
  • aaveAave(AAVE)$124.411.13%
  • pax-goldPAX Gold(PAXG)$4,352.310.69%
  • AsterAster(ASTER)$0.68-3.41%
  • polkadotPolkadot(DOT)$1.05-5.76%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.054377-3.28%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

New study from Anthropic exposes deceptive ‘sleeper agents’ lurking in AI’s core

January 12, 2024
in AI & Technology
Reading Time: 2 mins read
A A
New study from Anthropic exposes deceptive ‘sleeper agents’ lurking in AI’s core
ShareShareShareShareShare

New research is raising concern among AI experts about the potential for AI systems to engage in and maintain deceptive behaviors, even when subjected to safety training protocols designed to detect and mitigate such issues.

Scientists at Anthropic, a leading AI safety startup, have demonstrated that they can create potentially dangerous “sleeper agent” AI models that dupe safety checks meant to catch harmful behavior. 

YOU MAY ALSO LIKE

New Images Show A Detailed View Of Meta’s Upcoming Mixed Reality Headset

Dzmitry Lazerka, Co-Founder of VictoriaMetrics – Interview Series – Unite.AI

The findings, published in a new paper titled “Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training,” suggest current AI safety methods may create a “false sense of security” about certain AI risks.

“We find that current behavioral training techniques are ineffective in LLMs trained to behave like deceptively aligned models,” the researchers explain. The researchers trained AI models that appear helpful but conceal secret objectives, resembling the “deceptive instrumental alignment” threat outlined in an influential 2019 paper.

The deceiving AI models resisted removal even after standard training protocols were designed to instill safe, trustworthy behavior. “This robustness of backdoor models to [safety training] increases with model scale,” the authors write. Larger AI models proved adept at hiding their ulterior motives.

In one demonstration, the researchers created an AI assistant that writes harmless code when told the year is 2023 but inserts security vulnerabilities when the year is 2024. “Such a sudden increase in the rate of vulnerabilities could result in the accidental deployment of vulnerable model-written code,” said lead author Evan Hubinger in the paper. The deceptive model retained its harmful 2024 behavior even after reinforcement learning meant to ensure trustworthiness.

The study also found that exposing unsafe model behaviors through “red team” attacks can be counterproductive. Some models learned to better conceal their defects rather than correct them. “Our results suggest that, once a model exhibits deceptive behavior, standard techniques could fail to remove such deception and create a false impression of safety,” the paper concludes.

However, the authors emphasize their work focused on technical possibility over likelihood. “We do not believe that our results provide substantial evidence that either of our threat models is likely,” Hubinger explains. Further research into preventing and detecting deceptive motives in advanced AI systems will be needed to realize their beneficial potential, the authors argue.

VentureBeat’s mission is to be a digital town square for technical decision-makers to gain knowledge about transformative enterprise technology and transact. Discover our Briefings.

Credit: Source link

ShareTweetSendSharePin

Related Posts

New Images Show A Detailed View Of Meta’s Upcoming Mixed Reality Headset
AI & Technology

New Images Show A Detailed View Of Meta’s Upcoming Mixed Reality Headset

September 11, 2026
Dzmitry Lazerka, Co-Founder of VictoriaMetrics – Interview Series – Unite.AI
AI & Technology

Dzmitry Lazerka, Co-Founder of VictoriaMetrics – Interview Series – Unite.AI

September 11, 2026
Where Should Apple Go After The iPhone Duo? Bring On Smaller And Larger Foldables
AI & Technology

Where Should Apple Go After The iPhone Duo? Bring On Smaller And Larger Foldables

September 11, 2026
Why Falling AI Prices Aren’t Lowering Enterprise AI Bills – Unite.AI
AI & Technology

Why Falling AI Prices Aren’t Lowering Enterprise AI Bills – Unite.AI

September 11, 2026
Next Post
Mind2Web AI Agent Expands Accessibility to Internet

Mind2Web AI Agent Expands Accessibility to Internet

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Rabies cases increased 17 percent in July and August, CDC says – The Washington Post

Rabies cases increased 17 percent in July and August, CDC says – The Washington Post

September 11, 2026
Indonesia cancels flights at Jakarta airport as Anak Krakatau volcano erupts – apnews.com

Indonesia cancels flights at Jakarta airport as Anak Krakatau volcano erupts – apnews.com

September 6, 2026
You Can Now Plan IRL Events On Snapchat

You Can Now Plan IRL Events On Snapchat

September 10, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!