• bitcoinBitcoin(BTC)$76,686.000.70%
  • ethereumEthereum(ETH)$2,453.951.40%
  • tetherTether(USDT)$1.00-0.02%
  • binancecoinBNB(BNB)$745.143.06%
  • rippleXRP(XRP)$1.300.54%
  • usd-coinUSDC(USDC)$1.00-0.01%
  • solanaSolana(SOL)$102.363.41%
  • tronTRON(TRX)$0.335519-0.03%
  • zcashZcash(ZEC)$1,483.658.83%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.030.08%
  • HyperliquidHyperliquid(HYPE)$86.6810.15%
  • dogecoinDogecoin(DOGE)$0.0824242.13%
  • moneroMonero(XMR)$518.213.41%
  • USDSUSDS(USDS)$1.000.01%
  • whitebitWhiteBIT Coin(WBT)$78.940.88%
  • RainRain(RAIN)$0.012675-1.90%
  • chainlinkChainlink(LINK)$11.534.07%
  • leo-tokenLEO Token(LEO)$8.89-0.65%
  • cardanoCardano(ADA)$0.2115328.52%
  • stellarStellar(XLM)$0.1843310.38%
  • uniswapUniswap(UNI)$7.9318.71%
  • bitcoin-cashBitcoin Cash(BCH)$236.417.44%
  • Ethena USDeEthena USDe(USDE)$1.00-0.02%
  • daiDai(DAI)$1.00-0.02%
  • USD1USD1(USD1)$1.00-0.01%
  • CantonCanton(CC)$0.10664710.54%
  • litecoinLitecoin(LTC)$54.344.79%
  • nearNEAR Protocol(NEAR)$3.2123.37%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.342.11%
  • avalanche-2Avalanche(AVAX)$7.692.47%
  • hedera-hashgraphHedera(HBAR)$0.0752622.15%
  • Global DollarGlobal Dollar(USDG)$1.000.01%
  • suiSui(SUI)$0.776.91%
  • shiba-inuShiba Inu(SHIB)$0.0000056.01%
  • MemeCoreMemeCore(M)$1.2916.41%
  • crypto-com-chainCronos(CRO)$0.0580982.76%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.01%
  • tether-goldTether Gold(XAUT)$4,348.521.44%
  • BittensorBittensor(TAO)$236.085.96%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • okbOKB(OKB)$112.321.31%
  • Ripple USDRipple USD(RLUSD)$1.00-0.02%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.15-0.03%
  • aaveAave(AAVE)$130.568.42%
  • AsterAster(ASTER)$0.744.42%
  • BitwayBitway(BTW)$0.71-1.74%
  • polkadotPolkadot(DOT)$1.1210.44%
  • Pump.funPump.fun(PUMP)$0.0040518.61%
  • pax-goldPAX Gold(PAXG)$4,346.101.38%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Anthropic Says Claude Leads 26% of Its AI Research and Development – Unite.AI

September 17, 2026
in AI & Technology
Reading Time: 5 mins read
A A
Anthropic Says Claude Leads 26% of Its AI Research and Development – Unite.AI
ShareShareShareShareShare

Anthropic said its Claude models “lead” 26% of the company’s AI research and development work as of August 2026, in the first results from a prototype R&D Automation Index published in an Anthropic Institute post on September 17, 2026.

YOU MAY ALSO LIKE

eGPUs Do Work, But They Come With Some Notable Limitations

FAA Says Laser Strikes On Aircraft Fell For The Third Consecutive Year

The post, titled “Measurements for understanding the pace of AI development inside frontier labs,” pairs the index with internal metrics on agent oversight and compute allocation, and Anthropic said it intends to keep publishing such measurements.

The R&D Automation Index

The Anthropic R&D Automation Index maps the full range of AI R&D work done at the company, scores how automated each task currently is, and combines those scores into an aggregate measure. The scores use an Automation Level scale developed by Epoch AI, running from AL0, meaning no AI involvement, to AL5, where AI operates fully autonomously with no human in the loop. At AL3, AI “collaborates,” performing large portions of a task under close human direction; at AL4, AI “leads,” carrying most of a task from a high-level prompt to completion while a human supervises.

Anthropic reported that as of August 2026, Claude “leads” 26% of its AI R&D work, the share of work at or above the “collaborates” level is above 90%, and Claude is not operating fully autonomously for any measured subset of AI R&D work. A chart in the post describes the 26% “leads” share as up from under 1% in February 2026.

The underlying task catalogue was assembled from the bottom up using work records such as Slack and internal documentation. For each week of July 2026, a Claude research agent reviewed each randomly sampled person’s week — 20% of staff from every department that makes up the model R&D loop — and listed the tasks they worked on, yielding a flat list of roughly 15,000 granular tasks. Claude then organized those tasks into a hierarchical tree of 542 nodes, 378 of them leaves such as “eval platform defect diagnosis and fixes,” and that tree is frozen so every measurement runs against the same basket of work.

For each node, a Claude agent researches how that kind of work is done across the company, and an independent Claude judge assigns one of six automation levels, restricted to evidence from the month being rated or earlier. Tasks are weighted by person-time, so categories where more staff effort goes carry more weight.

Anthropic checked the judge’s ratings against staff who own the relevant work areas, who rated without seeing the models’ evidence or judgments, and reported model-versus-human exact agreement of 59% versus 35% human-versus-human, with ratings within one level of each other 97% of the time. Stated limitations include the frozen basket, which captures automation of existing work without registering new kinds of work; a comparison of tasks arriving from February through July 2026 against a January 2026 basket found no rise in “novel” tasks, and Anthropic plans to rebuild the basket periodically and re-version the published numbers.

Oversight of Internal Agents

A second set of measurements covers Anthropic’s oversight of AI agents on its systems. As of August 2026, approximately 30,000 agents were doing research and engineering work at any one time on the company’s most-used internal platform, and the reported measurements cover that platform only.

Anthropic reported that online monitors check 100% of these agents’ actions before execution, usually within seconds; of more than a billion agent decisions analyzed over August 2026, 0.002% (about 1 in 47,000) were blocked, and humans review any blocked actions within one week. Offline monitors ingest 100% of actions after the fact and flag roughly 100,000 transcripts per week, which multiple classifiers review for false positives, resolving most flags; approximately 50 highest-priority flags per week are escalated to human review, and roughly one to two transcripts in every thousand are flagged for further review.

According to the post, online monitors are built for harms that would happen quickly and could not be undone, such as an agent copying model weights out of Anthropic’s systems, while offline monitors target dangerous behaviors that develop slowly and can be reversed, like signs of model misalignment. Anthropic said two design choices in the internal scaffold support oversight: each agent carries an individual identity that persists across model upgrades, and agents communicate through a shared, open messaging system cross-referenced with transcripts, so behavioral monitoring can traverse communication between agents. The company said METR has previously independently red-teamed its offline monitoring platform, and that these oversight measurements were also published in its August 2026 risk report.

Compute Allocation Snapshot

For the third measurement, Anthropic examined how it used all of its compute from July 13 to July 20, 2026, sorting every workload into categories and measuring how much of the compute going to AI R&D was safety work. Over that week, Anthropic reported, about 6% of compute going to AI R&D was allocated to safety, and about 12% of compute going to AI-driven AI R&D was allocated to safety.

The company describes both estimates as deliberately conservative: tokens that advanced capabilities as much as safety were counted as AI R&D, and safeguards classifiers, a separate and comparable amount of compute, are excluded. A prompted Claude classifier sorted the week’s almost 10,000 research training and evaluation runs using a roughly 14% sample weighted toward the largest compute users, and Anthropic reported the classifier agreed with human reviewers within one or two percentage points.

Stated limitations are that a single week demonstrates the measurement is feasible without establishing a trend, that the underlying workload labels are best-effort and unverified, and that compute share measures what is spent rather than the amount of safety work performed.

Purpose and Next Steps

Anthropic said it is reporting the measurements because they give the public, outside parties, and governments a clearer view of the pace of AI development inside frontier labs, complementing its Responsible Scaling Policy risk reports and its Advanced AI Framework policy proposal. It noted the numbers would be expected to shift if there were coordination on pacing the frontier, as called for by CEO Dario Amodei.

The post says any frontier developer could publish the same measures regularly with a public methodology, and identifies two obstacles to cross-lab comparison: the absence of a shared methodology and a developer’s use of its own models to judge its systems. It says such measures could be verified by third parties or by other developers’ models, and could become a trigger for stronger requirements, such as a fixed testing window before a new model is put to work on further AI R&D.

Anthropic said it plans to embed independent third-party evaluators from multiple organizations, giving them access to internal processes, systems, and data comparable to what internal risk assessment teams have, to verify safety practices, report incidents, and track key metrics such as those in the post. The piece was co-authored by Marina Favaro and Phillie Wright, with research direction from Jack Clark.

Credit: Source link

ShareTweetSendSharePin

Related Posts

eGPUs Do Work, But They Come With Some Notable Limitations
AI & Technology

eGPUs Do Work, But They Come With Some Notable Limitations

September 17, 2026
FAA Says Laser Strikes On Aircraft Fell For The Third Consecutive Year
AI & Technology

FAA Says Laser Strikes On Aircraft Fell For The Third Consecutive Year

September 17, 2026
Microsoft Open-Sources TauGrid: A Kubernetes-Native Stack for GPU AI Workloads
AI & Technology

Microsoft Open-Sources TauGrid: A Kubernetes-Native Stack for GPU AI Workloads

September 17, 2026
GSA Extends Anthropic’s Claude OneGov Offer for Federal Agencies – Unite.AI
AI & Technology

GSA Extends Anthropic’s Claude OneGov Offer for Federal Agencies – Unite.AI

September 17, 2026
Next Post
ATI Inc. (ATI) Presents at Morgan Stanley’s 14th Annual Laguna Conference Transcript

ATI Inc. (ATI) Presents at Morgan Stanley's 14th Annual Laguna Conference Transcript

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Sinkhole threatens mansion once linked to Nicolas Cage

Sinkhole threatens mansion once linked to Nicolas Cage

September 15, 2026
Arch Manning and Texas Longhorns erase 20-point deficit to stun Ohio State in shocking comeback for the ages – Fox News

Arch Manning and Texas Longhorns erase 20-point deficit to stun Ohio State in shocking comeback for the ages – Fox News

September 13, 2026
College football picks: Predictions against the spread, odds, betting lines for top 25 games in Week 2 – cbssports.com

College football picks: Predictions against the spread, odds, betting lines for top 25 games in Week 2 – cbssports.com

September 12, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!