• bitcoinBitcoin(BTC)$64,843.000.90%
  • ethereumEthereum(ETH)$1,913.160.50%
  • tetherTether(USDT)$1.000.00%
  • binancecoinBNB(BNB)$590.960.10%
  • usd-coinUSDC(USDC)$1.000.00%
  • rippleXRP(XRP)$1.02-1.40%
  • solanaSolana(SOL)$73.611.40%
  • tronTRON(TRX)$0.3276500.10%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.032.30%
  • HyperliquidHyperliquid(HYPE)$54.09-3.60%
  • dogecoinDogecoin(DOGE)$0.0696951.00%
  • USDSUSDS(USDS)$1.000.00%
  • RainRain(RAIN)$0.0127141.20%
  • leo-tokenLEO Token(LEO)$9.760.00%
  • zcashZcash(ZEC)$508.213.30%
  • cardanoCardano(ADA)$0.1997830.10%
  • moneroMonero(XMR)$379.652.40%
  • whitebitWhiteBIT Coin(WBT)$56.080.70%
  • chainlinkChainlink(LINK)$8.170.00%
  • stellarStellar(XLM)$0.161250-0.20%
  • daiDai(DAI)$1.000.00%
  • bitcoin-cashBitcoin Cash(BCH)$214.581.00%
  • USD1USD1(USD1)$1.000.00%
  • Ethena USDeEthena USDe(USDE)$1.000.00%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.35-1.40%
  • CantonCanton(CC)$0.090902-0.50%
  • litecoinLitecoin(LTC)$45.500.30%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • Circle USYCCircle USYC(USYC)$1.130.00%
  • hedera-hashgraphHedera(HBAR)$0.068074-0.30%
  • paypal-usdPayPal USD(PYUSD)$1.000.00%
  • avalanche-2Avalanche(AVAX)$6.43-0.30%
  • suiSui(SUI)$0.67-0.10%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • shiba-inuShiba Inu(SHIB)$0.000005-1.10%
  • tether-goldTether Gold(XAUT)$4,330.632.60%
  • uniswapUniswap(UNI)$3.98-0.60%
  • crypto-com-chainCronos(CRO)$0.050472-5.40%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.140.00%
  • nearNEAR Protocol(NEAR)$1.59-4.30%
  • pax-goldPAX Gold(PAXG)$4,343.942.50%
  • okbOKB(OKB)$90.185.40%
  • BittensorBittensor(TAO)$193.320.60%
  • OndoOndo(ONDO)$0.346626-3.30%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.051410-1.30%
  • AsterAster(ASTER)$0.600.20%
  • HTX DAOHTX DAO(HTX)$0.000002-0.30%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • usddUSDD(USDD)$1.000.00%
  • MemeCoreMemeCore(M)$1.152.10%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Microsoft Open Sources code-testing-generator: a Polyglot Unit-Test Agent That Hits 92.1% Task Completion Versus 78.9% for Stock Copilot

August 7, 2026
in AI & Technology
Reading Time: 15 mins read
A A
Microsoft Open Sources code-testing-generator: a Polyglot Unit-Test Agent That Hits 92.1% Task Completion Versus 78.9% for Stock Copilot
ShareShareShareShareShare

Microsoft has open sourced code-testing-generator, a polyglot agent that writes unit tests and then proves they work. It ships in the dotnet-test plugin inside the MIT-licensed dotnet/skills repository.

The agent targets a gap that coding assistants usually leave open. A prompt like ‘generate unit tests’ does not say which framework, file location or assertions to use. code-testing-generator settles those decisions by reading the repository before it writes anything. It then plans, writes, runs and checks the tests it produces. On Microsoft’s internal 152-task benchmark, it completed 140 tasks against 120 for stock GitHub Copilot. Both setups used the same model and prompts.

Is it deployable

Yes. It is an agent definition along with skills, not a hosted service, so it runs inside your existing coding agent and code stays local.

  • Company stage: viable from solo maintainers upward. Startups and mid-market teams gain most, because the agent supplies repository research a small team has no time to encode. Enterprises can fork the language guidance to match internal frameworks.
  • Industries: regulated or audit-heavy software estates — financial services, healthcare, insurance, public sector — plus platform teams paying down legacy test debt.
  • Applications: backfilling tests on untested modules, generating tests for a pull-request diff, raising coverage before a release gate, and standardising conventions across polyglot monorepos.

What the agent actually does

It coordinates work through a Research-Plan-Implement (RPI) pipeline. It searches the repository for code needing tests, detects the language and test framework, reads existing tests for conventions, and finds the real build and test commands. That last step targets a specific failure: a test project that builds locally but never runs in CI because nothing registered it.

The agent then picks one of three strategies. Direct writes and validates tests immediately. Single pass runs one cycle. Iterative repeats it for large scopes or coverage targets. It never modifies production code, and avoids tests that call external URLs, bind ports or depend on timing.

The verification gate

Before reporting completion, the agent runs five checks. It reasons about small code changes that should make the tests fail, a lightweight form of mutation testing. It looks for weak or missing assertions. It maps every requested scenario to a test. It builds the full workspace and runs the full suite. It confirms the repository’s own test command discovers the new tests.

Benchmark results

On Microsoft’s internal benchmark of 152 tasks from real repositories, the agent completed 140 (92.1%) versus 120 (78.9%) for stock GitHub Copilot on the same model and prompts (63% fewer failures).

The gain is concentrated. On 89 vague prompts, the agent resolved 79 (88.8%) against 59 (66.3%), cutting failures from 30 to 10. On 63 detailed prompts, both scored 61 (96.8%). On 15 tasks targeting a specific diff, the agent passed all 15 and stock Copilot passed none.

Notably, the agent generated 2.3% fewer tests (6,963 vs 7,129) at effectively identical line coverage (72.4% vs 72.2%). Average task time was 359 seconds against 380. Token use per completed task was 3.2% higher.

On 45 .NET tasks, Claude Opus 4.8 reached 43/45 with the agent versus 35/45 stock; GPT-5.5 reached 41/45 versus 36/45. On the harder external SWE Atlas benchmark, completion was 16/44 versus 12/44.

Explainer: how the agent turns one prompt into verified tests

Key Takeaways

  • Open source, MIT-licensed, polyglot unit-test agent from Microsoft’s .NET team.
  • Research-Plan-Implement pipeline replaces one-shot generation with repository-aware planning.
  • 92.1% vs 78.9% task completion against stock Copilot on the same model.
  • Gains come almost entirely from vague prompts and diff-targeted requests.
  • Fewer tests, same coverage, 5.5% faster — reliability, not volume.

Check out the Technical details and Repo. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well.

Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.? Connect with us

YOU MAY ALSO LIKE

Apple’s Stumble, Amazon’s Surge and Anthropic’s Hacks | Bloomberg Tech 7/31/2026

AWS CEO Says AI Business Is ‘Just Massive’


Michal Sutter is a data science professional with a Master of Science in Data Science from the University of Padova. With a solid foundation in statistical analysis, machine learning, and data engineering, Michal excels at transforming complex datasets into actionable insights.

Credit: Source link

ShareTweetSendSharePin

Related Posts

Apple’s Stumble, Amazon’s Surge and Anthropic’s Hacks | Bloomberg Tech 7/31/2026
AI & Technology

Apple’s Stumble, Amazon’s Surge and Anthropic’s Hacks | Bloomberg Tech 7/31/2026

August 8, 2026
AWS CEO Says AI Business Is ‘Just Massive’
AI & Technology

AWS CEO Says AI Business Is ‘Just Massive’

August 8, 2026
AWS Is Investing to Keep Up With Demand, CEO Says
AI & Technology

AWS Is Investing to Keep Up With Demand, CEO Says

August 7, 2026
Valar Atomics Raises  Billion to Power the AI Era
AI & Technology

Valar Atomics Raises $1 Billion to Power the AI Era

August 7, 2026
Next Post
July 4 festivities delay start times due to extreme heat in Washington, D.C.

July 4 festivities delay start times due to extreme heat in Washington, D.C.

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Microsoft’s SkillOpt Shows Optimized Agent Skill Artifacts Transfer Across Model Scales and Between Codex and Claude Code Harnesses

Microsoft’s SkillOpt Shows Optimized Agent Skill Artifacts Transfer Across Model Scales and Between Codex and Claude Code Harnesses

August 6, 2026
Concentra Group Holdings Parent, Inc. (CON) Q2 2026 Earnings Call Transcript

Concentra Group Holdings Parent, Inc. (CON) Q2 2026 Earnings Call Transcript

August 7, 2026
Rep. Edwards drops competitive reelection bid after House censure recommendation – The Washington Post

Rep. Edwards drops competitive reelection bid after House censure recommendation – The Washington Post

August 6, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!