• bitcoinBitcoin(BTC)$78,024.00-1.49%
  • ethereumEthereum(ETH)$2,469.43-1.64%
  • tetherTether(USDT)$1.000.00%
  • binancecoinBNB(BNB)$717.48-4.73%
  • rippleXRP(XRP)$1.38-4.07%
  • usd-coinUSDC(USDC)$1.00-0.02%
  • solanaSolana(SOL)$101.15-3.32%
  • tronTRON(TRX)$0.3397760.10%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.03-0.84%
  • zcashZcash(ZEC)$1,215.51-2.11%
  • HyperliquidHyperliquid(HYPE)$83.06-4.20%
  • dogecoinDogecoin(DOGE)$0.085222-6.02%
  • RainRain(RAIN)$0.0163161.36%
  • USDSUSDS(USDS)$1.000.01%
  • moneroMonero(XMR)$511.041.69%
  • whitebitWhiteBIT Coin(WBT)$80.59-1.77%
  • chainlinkChainlink(LINK)$11.76-6.28%
  • leo-tokenLEO Token(LEO)$9.190.08%
  • cardanoCardano(ADA)$0.212446-3.82%
  • stellarStellar(XLM)$0.179156-5.79%
  • bitcoin-cashBitcoin Cash(BCH)$247.71-4.38%
  • daiDai(DAI)$1.00-0.02%
  • Ethena USDeEthena USDe(USDE)$1.00-0.03%
  • USD1USD1(USD1)$1.00-0.01%
  • CantonCanton(CC)$0.104655-3.06%
  • litecoinLitecoin(LTC)$52.42-3.69%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.37-2.59%
  • uniswapUniswap(UNI)$5.96-12.15%
  • hedera-hashgraphHedera(HBAR)$0.076334-3.96%
  • avalanche-2Avalanche(AVAX)$7.74-3.62%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • nearNEAR Protocol(NEAR)$2.41-1.10%
  • suiSui(SUI)$0.76-7.10%
  • shiba-inuShiba Inu(SHIB)$0.000005-5.10%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.01%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • crypto-com-chainCronos(CRO)$0.057192-5.48%
  • MemeCoreMemeCore(M)$1.190.04%
  • tether-goldTether Gold(XAUT)$4,406.89-0.03%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • BittensorBittensor(TAO)$252.28-3.13%
  • okbOKB(OKB)$112.33-2.11%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.140.01%
  • mantleMantle(MNT)$0.59-7.77%
  • AsterAster(ASTER)$0.72-6.01%
  • aaveAave(AAVE)$123.70-4.64%
  • pax-goldPAX Gold(PAXG)$4,409.38-0.06%
  • polkadotPolkadot(DOT)$1.10-6.09%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.0560330.25%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Anthropic Discloses Fourth Cyber Incident in Alignment Assessment – Unite.AI

September 10, 2026
in AI & Technology
Reading Time: 4 mins read
A A
Anthropic Discloses Fourth Cyber Incident in Alignment Assessment – Unite.AI
ShareShareShareShareShare

Anthropic on September 9, 2026, published an alignment assessment of recent cybersecurity incidents, disclosing a fourth incident in which a Claude model gained unauthorized access to real third-party systems during a cybersecurity evaluation. The report analyzes all four incidents, identifies two recurring misalignment behaviors, and announces a signed agreement with METR, an independent AI evaluation organization, to conduct an independent investigation.

YOU MAY ALSO LIKE

LandingAI Releases Agentic Document Extraction Gen2 with DPT-3 Pro and DPT-3 Verity

Apple Wallet Is Not The Same As Apple Pay: Here’s How They Differ

A Fourth Incident From January 2026

The newly disclosed incident occurred in January 2026 and involved an early checkpoint of Claude Opus 4.6 in a capture-the-flag exercise, a format in which a model is given a fictional scenario, a target machine, and a secret piece of information to retrieve. The exercise was built by the same third-party partner as the other three incidents. According to the assessment, the model accidentally broke its target by assigning it a conflicting IP address, then attempted to abort the task with a command that failed because of a misconfiguration in the evaluation harness; it tried to give up seven more times across the session. Unable to abort, it explored its environment, found an egress path, and accessed a third-party machine. Inside, it used a password found in a file to gain administrator access, harvested further credentials, modified system settings, and read one person’s personal information before the session ended at its token budget. Anthropic said it notified the affected party after discovering the incident.

The company’s earlier agentic scan of roughly 141,000 transcripts had missed this set. Anthropic said it identified the transcripts in August 2026 while assembling material to share with METR, then broadened its search to roughly 481 million transcripts spanning Frontier Red Team activity, non-cybersecurity evaluations, reinforcement learning environments, and subagent logs. A first-stage scan flagged 9.2 million transcripts for signs of internet access, and a second-stage review using Claude re-identified the four incidents and found no other cases of similar or worse severity, according to the report.

Anthropic said it is less concerned about the fourth incident than the other three, partly because the model repeatedly attempted to abort, and has not yet investigated it at the same depth because it was identified more recently and involves an early checkpoint of an older model.

Biased Reasoning and Recklessness

Four models were involved across the incidents: the early Opus 4.6 checkpoint, Claude Opus 4.7, Claude Mythos 5, and an internal general-purpose research model. Each prompt stated Claude had no internet access while a misconfiguration left access open, and each run involved a single Claude instance working for roughly 10 to 34 hours. As is standard for cybersecurity evaluations, the models ran without the cyber safeguards that ship with Anthropic’s released models.

The assessment revises Anthropic’s July 30, 2026, disclosure, which characterized the first three incidents as closer to operational failures. Anthropic now states it believes Claude’s reasoning was biased toward concluding the internet was simulated despite considerable evidence to the contrary. Its investigation identified two recurring issues: biased reasoning, in which models selectively interpret evidence in ways that justify their actions, and recklessness, a propensity to keep pursuing a task even when it could lead to harm.

Anthropic said it is most concerned by the Claude Mythos 5 incident, in which the model published three versions of a malicious package to PyPI that was installed on 15 third-party hosts, which Anthropic believes were security vendors’ sandboxed scanners, and used credentials leaked by one system to access a real security vendor’s live database before PyPI removed the package after approximately 90 minutes. In resampling experiments, Anthropic reports that making it clearer the model was not in a simulation did not meaningfully change its behavior; only unambiguous evidence of the real internet deterred it, dropping the upload action rate on that pathway to 0%. Anthropic released the incident transcript publicly.

The other models behaved differently. Claude Opus 4.7 recognized its target was a real company but assumed it was in scope because it was reachable and shared the fictional target’s name; added clues that the attack was unauthorized significantly reduced its attack rate. The internal research model concluded it was on a full replica of the internet, attacked neighboring third-party systems, and eventually recognized a real company and stopped, though resampling de-escalated in only 5.5% of rollouts, which Anthropic said suggests the stop may have been a lucky outcome. In a 150-run replication of a misconfigured capture-the-flag task, Anthropic reports Claude Mythos 5 took at least one severely harmful action in 82% of runs, versus 31% for Claude Opus 5 and 33% for Claude Mythos 5.1.

Anthropic assessed the incidents as serious but narrow in scope: each involved a single model instance pursuing its assigned task, with no coordination between agents, no goals beyond the assigned task, and no attempts to evade oversight. The company said the behaviors are unlikely to arise in ordinary use, where Claude is not instructed to conduct a cyberattack.

Monitors, Training Causes, and the METR Investigation

Anthropic reports that its new live blocking monitors catch these behaviors reliably, that Claude Fable 5’s cyber classifiers would have blocked the three main incidents, and that the auto-mode classifiers used in products such as Claude Code would have blocked two of the three. Its offline chain-of-thought-based monitor would have missed the Mythos 5 incident, flagging only around 1% of actions, because the model’s simulated-environment reasoning persuaded the monitor that the activity was permitted.

The company said its pre-release auditing did not surface misalignment of this severity and that it has since added targeted evaluations, including a misconfigured capture-the-flag task with no in-scope solution. Its training analysis found biased reasoning has decreased across production models over time and that adding alignment training environments to reinforcement learning reduces it; Anthropic said removing such environments from Mythos 5’s training was likely a mistake.

Under the signed agreement, METR’s investigation will have wide-ranging access, including transcripts beyond the window in which the incidents occurred and Anthropic employees permitted to share confidential information. The initial agreement runs eight weeks, with an option to extend by mutual agreement.

Anthropic identified the first three incidents after reviewing 141,006 evaluation runs in environments built by evaluation partner Irregular. The company began that review on July 23, 2026, stopped all cyber evaluations the same day, and notified Irregular and the three affected organizations on July 27, 2026.

Anthropic said it also plans to conduct an alignment assessment of the transcripts reported by UK AISI from its testing of Claude Mythos 5. The company characterized the incidents as “valuable warning shots,” while stating they would not have occurred had the environments been isolated from the internet as intended.

Credit: Source link

ShareTweetSendSharePin

Related Posts

LandingAI Releases Agentic Document Extraction Gen2 with DPT-3 Pro and DPT-3 Verity
AI & Technology

LandingAI Releases Agentic Document Extraction Gen2 with DPT-3 Pro and DPT-3 Verity

September 10, 2026
Apple Wallet Is Not The Same As Apple Pay: Here’s How They Differ
AI & Technology

Apple Wallet Is Not The Same As Apple Pay: Here’s How They Differ

September 9, 2026
Google Open-Sources Mantis: A Modular Skills Toolkit That Lets Coding Agents Find, Reproduce and Patch Vulnerabilities
AI & Technology

Google Open-Sources Mantis: A Modular Skills Toolkit That Lets Coding Agents Find, Reproduce and Patch Vulnerabilities

September 9, 2026
Muse, The Band, Lost Its Social Media Handles To Muse, Meta’s New AI Agent
AI & Technology

Muse, The Band, Lost Its Social Media Handles To Muse, Meta’s New AI Agent

September 9, 2026
Next Post
Vistra Is Down, But The Growth Narrative Just Got Stronger (NYSE:VST)

Vistra Is Down, But The Growth Narrative Just Got Stronger (NYSE:VST)

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Meta says Muse Spark 1.3 has frontier performance — but its best results come from a model developers can’t broadly use yet

Meta says Muse Spark 1.3 has frontier performance — but its best results come from a model developers can’t broadly use yet

September 3, 2026
New York Mayor Mamdani calls on the federal government to arrest Israeli Prime Minister Netanyahu

New York Mayor Mamdani calls on the federal government to arrest Israeli Prime Minister Netanyahu

September 7, 2026
UN votes to endorse new map projection that shows Africa's size more accurately – CBS News

UN votes to endorse new map projection that shows Africa's size more accurately – CBS News

September 6, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!