• bitcoinBitcoin(BTC)$86,285.003.36%
  • ethereumEthereum(ETH)$2,755.222.66%
  • tetherTether(USDT)$1.000.01%
  • binancecoinBNB(BNB)$778.121.67%
  • rippleXRP(XRP)$1.543.57%
  • usd-coinUSDC(USDC)$1.000.00%
  • solanaSolana(SOL)$122.103.67%
  • tronTRON(TRX)$0.334413-0.84%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.02-0.47%
  • zcashZcash(ZEC)$1,382.64-1.80%
  • HyperliquidHyperliquid(HYPE)$90.231.48%
  • dogecoinDogecoin(DOGE)$0.0972433.12%
  • chainlinkChainlink(LINK)$14.391.27%
  • moneroMonero(XMR)$546.37-0.02%
  • whitebitWhiteBIT Coin(WBT)$86.093.25%
  • USDSUSDS(USDS)$1.000.01%
  • cardanoCardano(ADA)$0.2551603.73%
  • RainRain(RAIN)$0.012001-1.95%
  • leo-tokenLEO Token(LEO)$8.951.16%
  • stellarStellar(XLM)$0.2249460.85%
  • nearNEAR Protocol(NEAR)$4.97-4.80%
  • bitcoin-cashBitcoin Cash(BCH)$315.752.82%
  • uniswapUniswap(UNI)$9.080.81%
  • litecoinLitecoin(LTC)$70.424.93%
  • avalanche-2Avalanche(AVAX)$11.100.23%
  • Ethena USDeEthena USDe(USDE)$1.000.00%
  • suiSui(SUI)$1.192.80%
  • CantonCanton(CC)$0.120401-2.21%
  • Blockchain USDBlockchain USD(USDB)$0.871,000.00%
  • hedera-hashgraphHedera(HBAR)$0.104896-0.14%
  • daiDai(DAI)$1.00-0.02%
  • USD1USD1(USD1)$1.000.03%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.553.29%
  • BitwayBitway(BTW)$1.453.13%
  • BittensorBittensor(TAO)$312.641.35%
  • shiba-inuShiba Inu(SHIB)$0.0000063.70%
  • crypto-com-chainCronos(CRO)$0.0691340.50%
  • quant-networkQuant(QNT)$236.26-21.12%
  • tether-goldTether Gold(XAUT)$4,178.990.58%
  • Global DollarGlobal Dollar(USDG)$1.00-0.01%
  • paypal-usdPayPal USD(PYUSD)$1.000.02%
  • aaveAave(AAVE)$181.799.63%
  • Pump.funPump.fun(PUMP)$0.0059187.22%
  • okbOKB(OKB)$122.661.41%
  • Ripple USDRipple USD(RLUSD)$1.00-0.01%
  • EthenaEthena(ENA)$0.246502-5.37%
  • OndoOndo(ONDO)$0.510.81%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • MemeCoreMemeCore(M)$1.052.34%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.15-0.03%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Anthropic Adds Plugin Evals to Claude Code: 6 Grader Types, a No-Plugin Baseline, and a CI Gate for Skills

September 11, 2026
in AI & Technology
Reading Time: 18 mins read
A A
Anthropic Adds Plugin Evals to Claude Code: 6 Grader Types, a No-Plugin Baseline, and a CI Gate for Skills
ShareShareShareShareShare

Anthropic has published a new plugin evals workflow for Claude Code. The claude plugin eval command runs a plugin against realistic prompts, grades what Claude produced, and compares the result with a run where the plugin is not loaded. It answers 3 questions plugin developers could not previously measure: does the skill trigger, does it survive an edit or a new model, and does it beat a bare model.

Deployable: Yes. It runs on Claude Code v2.1.269 or later against any directory with a plugin.json or .claude-plugin/plugin.json manifest, or a skills-directory plugin. Every eval run and judge grader is a real model call billed to your plan or API account.

YOU MAY ALSO LIKE

Cloudflare Releases Clef and Clef-flash: Open-Weight Decision Models That Return Typed Probabilities Instead of Text

A Coding Guide to Google Research’s Kauldron: Configs That Are Plain Data, Components Wired by String, and a JAX Trainer You Can Read End to End

What a case looks like

An eval suite lives in an evals/ directory inside the plugin. Each case is a subdirectory holding a prompt.md and a graders/ folder. The prompt body goes to Claude exactly as written, and @path mentions are not expanded. Frontmatter on prompt.md can set max_turns (default 10), timeout_seconds (default 300), model, tags, and allowed_tools.

Graders are markdown files whose frontmatter sets a type, an optional weight, and an optional arm. There are 6 types. Four cost nothing because they are computed from the transcript and the files on disk: regex, tool_used, tool_order, and file_exists. Two call a judge model and add to the bill: llm, which scores the reply against prose criteria you write, and baseline, which compares it against a reference answer.

claude plugin eval init reads the plugin, asks what a good result looks like, proposes cases and graders, tries them, and writes the files. In CI, --bare writes a blank template instead.

The number that matters is Δ

By default every case runs twice: a with-arm where the plugin is loaded and a without-arm where it is not. Their difference, Δ, is what the plugin contributed. If a case scores 1.0 in both arms, the plugin is not why it passed. The docs example output shows a single case at WITH 1.00, W/OUT 0.33, Δ +0.67 across 6 runs, costing an estimated $0.41 and taking 74 seconds. A grader marked with-only, typically tool_used: Skill, is reported as an indicator and excluded from the score, since the without-arm has no skill to fire.

Anthropic calls out the most common first finding: a Δ near zero with the tool_used: Skill grader failing, which means Claude is not choosing the skill on natural phrasing. That is the defect claude plugin validate cannot see, because it checks manifest syntax and schema rather than behavior.

Results land under evals/results//report.html with per-grader verdicts and judge votes. Where the account supports it, the report is also published to claude.ai unless --no-publish is set.

Cost and CI

A suite makes roughly cases × runs × arms agent runs, plus 3 short judge calls per llm or baseline grader per run, and results vary between runs. The documented CI invocation is:

claude plugin eval . \
  --trust-plugin \
  --json results.json \
  --threshold 0.8 \
  --model claude-sonnet-5 \
  --judge-model claude-haiku-4-5 \
  --no-publish \
  --max-cost-usd 20

The runner needs a Claude Code install and credentials such as ANTHROPIC_API_KEY. Without --trust-plugin, an untrusted checkout is refused with exit 1 when there is no terminal. Report problems never change the exit code, and --json suppresses progress output.

Interactive explainer

Credit: Source link

ShareTweetSendSharePin

Related Posts

Cloudflare Releases Clef and Clef-flash: Open-Weight Decision Models That Return Typed Probabilities Instead of Text
AI & Technology

Cloudflare Releases Clef and Clef-flash: Open-Weight Decision Models That Return Typed Probabilities Instead of Text

October 2, 2026
A Coding Guide to Google Research’s Kauldron: Configs That Are Plain Data, Components Wired by String, and a JAX Trainer You Can Read End to End
AI & Technology

A Coding Guide to Google Research’s Kauldron: Configs That Are Plain Data, Components Wired by String, and a JAX Trainer You Can Read End to End

October 2, 2026
Florida County Finds 11 Unpermitted Flock Cameras With Unidentified Owners
AI & Technology

Florida County Finds 11 Unpermitted Flock Cameras With Unidentified Owners

October 1, 2026
Suno Launches Speech Beta, Pairing Voice and Music in One Model – Unite.AI
AI & Technology

Suno Launches Speech Beta, Pairing Voice and Music in One Model – Unite.AI

October 1, 2026
Next Post
I Made This Jelly Game With GPT-6 Astra

I Made This Jelly Game With GPT-6 Astra

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Porsche: How To Handle Automotive Drawdown (And Potential Upside)

Porsche: How To Handle Automotive Drawdown (And Potential Upside)

September 28, 2026
Angie Nixon says Florida ‘is not red, it just hasn’t been invested in by Democrats’: Full interview

Angie Nixon says Florida ‘is not red, it just hasn’t been invested in by Democrats’: Full interview

September 26, 2026
Exa Launches Agent Ultra: A Subagent Swarm Deep Research API Built for Exhaustive List Building

Exa Launches Agent Ultra: A Subagent Swarm Deep Research API Built for Exhaustive List Building

September 26, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!