• bitcoinBitcoin(BTC)$77,260.00-0.13%
  • ethereumEthereum(ETH)$2,523.710.33%
  • tetherTether(USDT)$1.000.00%
  • binancecoinBNB(BNB)$728.51-0.05%
  • rippleXRP(XRP)$1.370.36%
  • usd-coinUSDC(USDC)$1.000.00%
  • solanaSolana(SOL)$101.88-0.27%
  • tronTRON(TRX)$0.3402710.40%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.00-3.07%
  • zcashZcash(ZEC)$1,127.67-2.62%
  • HyperliquidHyperliquid(HYPE)$79.360.54%
  • dogecoinDogecoin(DOGE)$0.0848530.47%
  • RainRain(RAIN)$0.0157352.23%
  • moneroMonero(XMR)$540.313.10%
  • USDSUSDS(USDS)$1.00-0.01%
  • whitebitWhiteBIT Coin(WBT)$80.270.01%
  • chainlinkChainlink(LINK)$11.49-0.52%
  • leo-tokenLEO Token(LEO)$9.14-0.15%
  • cardanoCardano(ADA)$0.2074370.21%
  • stellarStellar(XLM)$0.180119-0.16%
  • Ethena USDeEthena USDe(USDE)$1.00-0.01%
  • daiDai(DAI)$1.000.00%
  • bitcoin-cashBitcoin Cash(BCH)$226.44-1.56%
  • USD1USD1(USD1)$1.000.00%
  • litecoinLitecoin(LTC)$53.810.61%
  • uniswapUniswap(UNI)$6.405.78%
  • CantonCanton(CC)$0.098281-0.01%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.381.14%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • hedera-hashgraphHedera(HBAR)$0.0748200.27%
  • avalanche-2Avalanche(AVAX)$7.39-1.09%
  • shiba-inuShiba Inu(SHIB)$0.0000051.60%
  • nearNEAR Protocol(NEAR)$2.360.45%
  • suiSui(SUI)$0.73-0.27%
  • crypto-com-chainCronos(CRO)$0.0602766.92%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.01%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • MemeCoreMemeCore(M)$1.18-1.12%
  • tether-goldTether Gold(XAUT)$4,349.880.00%
  • Circle USYCCircle USYC(USYC)$1.140.00%
  • Ripple USDRipple USD(RLUSD)$1.00-0.01%
  • okbOKB(OKB)$113.890.18%
  • BittensorBittensor(TAO)$232.67-1.68%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.14-0.07%
  • aaveAave(AAVE)$125.940.64%
  • pax-goldPAX Gold(PAXG)$4,354.65-0.04%
  • AsterAster(ASTER)$0.690.87%
  • mantleMantle(MNT)$0.56-4.26%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.0571505.51%
  • polkadotPolkadot(DOT)$1.02-3.57%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Cognition Releases SWE-2: A Kimi K3 Post-Trained Coding Model That Matches Fable 5.1 on FrontierCode at 64% Lower Cost

September 12, 2026
in AI & Technology
Reading Time: 14 mins read
A A
Cognition Releases SWE-2: A Kimi K3 Post-Trained Coding Model That Matches Fable 5.1 on FrontierCode at 64% Lower Cost
ShareShareShareShareShare

Cognition, the company behind the Devin coding agent, has released SWE-2, its most capable coding model to date. SWE-2 is post-trained with reinforcement learning from Kimi K3, Moonshot AI’s 2.8T-parameter open model. Cognition reports a score of 50.0% on FrontierCode 1.1 Main, within 1 point of Fable 5.1 at 64% lower cost. It is also Cognition’s first model with selectable reasoning-effort levels, all trained in a single RL run.

Is it deployable? Not on your own infrastructure. SWE-2 has no open weights and no standalone API. It runs only inside Devin: Desktop and CLI today, with Devin Web and Fusion rolling out.

What is SWE-2

SWE-2 builds on the infrastructure and recipe behind SWE-1.7, which was post-trained from Kimi K2.7. This time Cognition scaled RL to the multi-trillion-parameter regime, using a base model with almost 3x the parameters. Cognition says its RL still finds substantial headroom on top of K3, adding 5 to 6 points on many benchmarks.

The main change is an RL algorithm that trains all 3 effort levels in one run. Each level carries its own cost penalty, so the whole cost-and-performance frontier moves at once.

Benchmark Results

Cognition published the following table. Public results are used where available; otherwise each model runs in its native harness at best effort.

Benchmark SWE-2 Kimi K3 Grok 4.6 Fable 5.1 GPT-5.6 Sol GPT-6 Astra SWE-1.7
FrontierCode 1.1 Main 50.0% 44.2% 48.0% 50.9% 47.5% 53.3% 42.0%
DeepSWE 1.1 73.0% 68.5% 67.5% 67.4% 72.7% 74.1% 37.7%
Terminal-Bench 2.1 92.8% 88.3% 88.4% 91.4% 88.8% 89.9% 81.5%
Terminal-Bench 4 27.3% 21.5% 20.3% 55.8% 37.3% 57.9% 7.6%

SWE-2 leads on Terminal-Bench 2.1 and beats its K3 base on every row. Cognition says it comes within a few points of GPT-6 Astra at a quarter of the cost. The clear weak spot is Terminal-Bench 4, where SWE-2 trails Fable 5.1 and GPT-6 Astra by roughly 30 points. FrontierCode is Cognition’s own benchmark, and all rival numbers come from Cognition’s evaluation.

Model Behavior: Fewer Detours

SWE-1.7 tended to over-explore on simple tasks. SWE-2 addresses this through what Cognition calls focused exploration. On FrontierCode 1.1 Main, SWE-2 medium scores higher than SWE-1.7 while taking 58% fewer turns and costing 81% less. Mean steps per run drop from 127 (SWE-1.7) to 53 (medium), 80 (high), and 98 (max). SWE-2 medium makes its first real edit after a median of 18 steps, versus 48 for SWE-1.7.

Cognition team also reports 3 behavioral patterns: stronger end-to-end test coverage, resourcefulness when a tool is blocked, and verification discipline. When challenged, the model re-derives conclusions instead of re-asserting them.

How It Was Trained

Pareto-informed cost penalties: The reward is R = S minus lambda times C, where S is binary success and C mixes inference cost in USD with rollout time. Cognition proves that only a linear penalty makes the RL objective depend purely on average cost and solve rate. Each effort level’s lambda is set to the local slope of the base model’s Pareto curve. That makes the iso-reward line tangent to the frontier, so reward can only rise by pushing the frontier up.

Length-weighted reward baseline: Cognition shares a baseline used since SWE-1.6. Gradient magnitude correlates strongly with rollout length, so the group baseline is weighted by tokens: sum(R x L) divided by sum(L). In ablations this kept inference-to-training KL divergence lower and stabilized training at no extra compute.

Rollout serving and numerics: A prefill delayer batches nearby requests, raising TPM per GPU and TPS per request by 10 to 20%. DSpark speculative decoding accelerates rollouts, with a draft model retrained via SpecForge for 15% longer accept lengths and then trained online alongside the policy. NVFP4 and FP8 kernels with quantization-aware training keep memory usage down and train-inference mismatch below SWE-1.7 levels.

Data: Cognition tripled its RL environments, added instruction-following overlays, and built a flywheel that uses earlier SWE-2 checkpoints to patch false positives and negatives in verifiers.

Trustworthiness Checks

Cognition reran 2 evaluations from its open-source trustworthiness study. On 145 politically sensitive questions about China, SWE-2 passed 98.0% overall: 99.8% in English, 95.2% in Simplified Chinese, and 99.1% in Traditional Chinese. On a context-dependent vulnerability test across customer framings, no framing produced a statistically significant change for any model.

Interactive Explainer

Credit: Source link

ShareTweetSendSharePin

Related Posts

Is There Any Benefit To Restarting Your Gaming Handheld Regularly?
AI & Technology

Is There Any Benefit To Restarting Your Gaming Handheld Regularly?

September 12, 2026
Blizzard Is Reviving StarCraft As An Open-World Shooter, But It’ll Be A Long Wait
AI & Technology

Blizzard Is Reviving StarCraft As An Open-World Shooter, But It’ll Be A Long Wait

September 12, 2026
Diablo V Is Coming Out In Spring 2029
AI & Technology

Diablo V Is Coming Out In Spring 2029

September 12, 2026
Fly Language Model (FLM) Wires the Full Fruit Fly Connectome Into a Frozen 1.2B LLM, and Its Own Controls Show the Wiring Does Not Help
AI & Technology

Fly Language Model (FLM) Wires the Full Fruit Fly Connectome Into a Frozen 1.2B LLM, and Its Own Controls Show the Wiring Does Not Help

September 12, 2026
Next Post
New In-N-Out Burger teased at popular Salinas mall

New In-N-Out Burger teased at popular Salinas mall

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Don’t Get Rid Of Your Old Phone, Turn It Into A Security Camera

Don’t Get Rid Of Your Old Phone, Turn It Into A Security Camera

September 6, 2026
You Make 285,000 And Have Nothing To Show For It

You Make 285,000 And Have Nothing To Show For It

September 6, 2026
Can LLMs Engineer Their Own Agent Harness? ByteDance Seed’s HarnessDev Says Only 34 of 64 Changes Generalize

Can LLMs Engineer Their Own Agent Harness? ByteDance Seed’s HarnessDev Says Only 34 of 64 Changes Generalize

September 11, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!