• bitcoinBitcoin(BTC)$66,331.001.63%
  • ethereumEthereum(ETH)$1,929.461.02%
  • tetherTether(USDT)$1.000.01%
  • binancecoinBNB(BNB)$572.630.05%
  • usd-coinUSDC(USDC)$1.000.00%
  • rippleXRP(XRP)$1.142.51%
  • solanaSolana(SOL)$78.210.39%
  • tronTRON(TRX)$0.3292071.03%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.010.47%
  • HyperliquidHyperliquid(HYPE)$60.80-2.56%
  • dogecoinDogecoin(DOGE)$0.0735371.65%
  • RainRain(RAIN)$0.01554010.21%
  • USDSUSDS(USDS)$1.000.00%
  • leo-tokenLEO Token(LEO)$9.72-0.09%
  • zcashZcash(ZEC)$532.14-1.73%
  • whitebitWhiteBIT Coin(WBT)$57.751.39%
  • moneroMonero(XMR)$355.784.48%
  • stellarStellar(XLM)$0.1930023.12%
  • cardanoCardano(ADA)$0.1748892.88%
  • chainlinkChainlink(LINK)$8.701.42%
  • CantonCanton(CC)$0.1271872.13%
  • daiDai(DAI)$1.000.01%
  • bitcoin-cashBitcoin Cash(BCH)$225.081.74%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.589.85%
  • USD1USD1(USD1)$1.000.00%
  • Ethena USDeEthena USDe(USDE)$1.000.00%
  • litecoinLitecoin(LTC)$46.61-1.12%
  • Global DollarGlobal Dollar(USDG)$1.000.08%
  • suiSui(SUI)$0.770.95%
  • hedera-hashgraphHedera(HBAR)$0.0704935.64%
  • Circle USYCCircle USYC(USYC)$1.13-0.01%
  • avalanche-2Avalanche(AVAX)$6.60-0.16%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.01%
  • crypto-com-chainCronos(CRO)$0.0585711.38%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • tether-goldTether Gold(XAUT)$4,116.812.10%
  • nearNEAR Protocol(NEAR)$1.93-2.97%
  • shiba-inuShiba Inu(SHIB)$0.0000040.15%
  • uniswapUniswap(UNI)$3.680.71%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.140.37%
  • OndoOndo(ONDO)$0.40139311.71%
  • BittensorBittensor(TAO)$201.183.84%
  • pax-goldPAX Gold(PAXG)$4,115.892.14%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.0571150.45%
  • okbOKB(OKB)$82.240.88%
  • AsterAster(ASTER)$0.63-0.88%
  • HTX DAOHTX DAO(HTX)$0.000002-0.11%
  • Ripple USDRipple USD(RLUSD)$1.000.02%
  • MemeCoreMemeCore(M)$1.18-0.80%
  • usddUSDD(USDD)$1.000.01%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Poolside Releases Laguna S 2.1, an Open-Weight Agentic Coding Model Punching Above Its Weight Class on SWE-Bench Multilingual

July 22, 2026
in AI & Technology
Reading Time: 40 mins read
A A
Poolside Releases Laguna S 2.1, an Open-Weight Agentic Coding Model Punching Above Its Weight Class on SWE-Bench Multilingual
ShareShareShareShareShare

Poolside has released Laguna S 2.1, a 118B-parameter open-weight model built for agentic coding. It is a Mixture-of-Experts (MoE) model with 8B activated parameters per token. It supports a context window of up to 1M tokens in both thinking and no-thinking modes. The weights are on Hugging Face under an OpenMDW-1.1 license, and the model is small enough to run on a single NVIDIA DGX Spark.

On long-horizon coding benchmarks, Laguna S 2.1 holds its own against models several times its size, including DeepSeek-V4-Pro-Max, NVIDIA’s Nemotron 3 Ultra, and Thinking Machines’ Inkling. Laguna S 2.1 is a scale-up of the Laguna XS family, trained on the same pre-training data as XS 2.1.

What is Laguna S 2.1

The model activates roughly 6.8% of its parameters on any given token. All 118B parameters remain resident in memory, but only ~8B route through the network per step. That sparsity is why a mid-size model can behave like a larger one while staying cheap to serve.

Poolside team publishes weights in BF16, FP8, INT4, and NVFP4, along with official GGUF and MLX conversions and DFlash draft models. It went from the start of training to launch in under nine weeks. Pre-training began on 22 May 2026 on 4,096 NVIDIA H200 GPUs. It is the first Poolside model where reinforcement learning ran in FP8 precision.

Performance

Laguna S 2.1 scores 70.2% on Terminal-Bench 2.1 with thinking enabled. That places it first among open, disclosed-size models on Poolside’s compiled leaderboard, behind only larger or closed systems. On SWE-Bench Multilingual it scores 78.5%, topping the published table outright. The full comparison Poolside released is below.

Benchmark Laguna S 2.1 (118B-A8B) Tencent Hy3 (295B-A21B) Inkling (975B-A41B) Nemotron 3 Ultra (550B-A55B) DeepSeek-V4-Pro-Max (1.6T-A49B) Kimi K3 (2.8T-A50B) Qwen 3.7 Max Muse Spark 1.1 Claude Fable 5
Terminal-Bench 2.1 70.2 71.7 63.8 56.4 64.0 88.3 74.5 80.0 88.0
SWE-Bench Multilingual 78.5 75.8 – 67.7 76.2 – 78.3 – –
SWE-Bench Pro (Public) 59.4 57.9 54.3 – 55.4 – 60.6 61.5 80.3
DeepSWE v1.1 40.4 – – – 9.0 69.0 – 53.3 70.0
SWE Atlas (Codebase QnA) 46.2 – – – 27.2 – – 42.2 –
Toolathlon Verified 49.7 – 45.5 34.3 55.9 – – 75.6 –

The clearest signal is DeepSWE v1.1, which still has real headroom. There, Laguna S 2.1 scores 40.4% against DeepSeek-V4-Pro-Max’s 9.0%, with roughly one-sixth the active parameters. Closed frontier models such as Claude Fable 5 and Kimi K3 still lead on several benchmarks. Poolside’s claim is about the weight class, not the outright top. Trajectories from the final evaluation set is published at trajectories.poolside.ai.

YOU MAY ALSO LIKE

Poolside drops Laguna S 2.1, an open-weight coding model that beats rivals 10x its size

France Will Ban Social Media For Children Under 15

Two thinking modes, and where the score comes from

Laguna S 2.1 has two modes: off and max, with max enabled by default. In max mode the model sets its own test-time compute budget. Poolside is shipping without user-configurable low/medium/high effort control for now.

Max thinking lifts Terminal-Bench 2.1 from 60.4% to 70.2%. It lifts DeepSWE from 16.5% to 40.4%. Those gains cost tokens: DeepSWE trajectories run about 249k completion tokens in thinking mode against 99k without. Poolside team reports coherent, productive reasoning running for hours and hundreds of thousands of tokens.

Three published runs

Poolside team shared three unedited trajectories to show behavior rather than scores. In one, the model built a working HTML/CSS browser engine from an empty folder. That run took 181 steps across a 50-minute session, with no human intervention, and it validated its own output against headless Chromium.

In a second, the model optimized Poolside’s own agent harness. It made the harness 5.2% faster with roughly 71% lower memory allocation, replacing O(n²) string concatenation with buffers. In a third, it re-derived Erdős Problem #397 offline in Perl over 68 minutes, since the sandbox had no Python. That result is an independent rediscovery; GPT-5.2 Pro solved the same problem in January 2026, and Laguna’s knowledge cutoff is November 2025.

Can you actually deploy it?

Sizing uses the full 118B parameters, not the 8B active count, because every expert stays in memory. At 4-bit (NVFP4 or INT4) the weights need about 59 GB, which fits comfortably on a single DGX Spark’s 128 GB of unified memory. At FP8 they need about 118 GB, still within a single Spark or a single H200. At BF16 they need about 236 GB, which calls for two linked Sparks or a multi-GPU datacenter node.

Poolside worked with NVIDIA to optimize inference from TRT-LLM serving to NVFP4 on Blackwell, down to a single DGX Spark. It shipped day-one support for vLLM, SGLang, and Ollama. Hosted access runs through OpenRouter, free at 256K context and paid at the full 1M context for $0.10 / $0.20 / $0.01 per 1M input / output / cache-read tokens. The model is also on Baseten, Kilo, Prime Intellect Prime Lab, and ZML.

Key Takeaways

  • Laguna S 2.1 is a 118B-total / 8B-active MoE coding model with a 1M-token context, open under OpenMDW-1.1.
  • It scores 70.2% on Terminal-Bench 2.1 and 78.5% on SWE-Bench Multilingual, leading open disclosed-size models.
  • At 4-bit it runs on a single NVIDIA DGX Spark; FP8 fits one Spark or H200, BF16 needs two.
  • Default ‘max thinking’ drives most of the score, at a real token cost (DeepSWE: 16.5% → 40.4%).
  • Trained in under nine weeks on 4,096 H200 GPUs, it is Poolside’s third model in under three months.

Sources: Poolside Technical— Introducing Laguna S 2.1, Robert McHardy on X, Poolside on X and Hugging Face model card.


Asif Razzaq is the CEO of Marktechpost Media Inc.. As a visionary entrepreneur and engineer, Asif is committed to harnessing the potential of Artificial Intelligence for social good. His most recent endeavor is the launch of an Artificial Intelligence Media Platform, Marktechpost, which stands out for its in-depth coverage of machine learning and deep learning news that is both technically sound and easily understandable by a wide audience. The platform boasts of over 2 million monthly views, illustrating its popularity among audiences.

Credit: Source link

ShareTweetSendSharePin

Related Posts

Poolside drops Laguna S 2.1, an open-weight coding model that beats rivals 10x its size
AI & Technology

Poolside drops Laguna S 2.1, an open-weight coding model that beats rivals 10x its size

July 21, 2026
France Will Ban Social Media For Children Under 15
AI & Technology

France Will Ban Social Media For Children Under 15

July 21, 2026
Stop adding more GPUs: Weka’s new storage platform reduces load by caching 100% of an AI model’s pre-calculated tokens
AI & Technology

Stop adding more GPUs: Weka’s new storage platform reduces load by caching 100% of an AI model’s pre-calculated tokens

July 21, 2026
Substack Is Adding An AI Detection Feature
AI & Technology

Substack Is Adding An AI Detection Feature

July 21, 2026
Next Post
Man accused of setting fire outside NYC federal building

Man accused of setting fire outside NYC federal building

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
NBA Summer League scouting report: How the top rookies already look stardom-bound – The New York Times

NBA Summer League scouting report: How the top rookies already look stardom-bound – The New York Times

July 17, 2026
Kimi K3 vs DeepSeek V4 Pro vs GLM-5.2: Open Trillion-Scale MoE Models Compared on Benchmarks, License, and Serving Cost

Kimi K3 vs DeepSeek V4 Pro vs GLM-5.2: Open Trillion-Scale MoE Models Compared on Benchmarks, License, and Serving Cost

July 19, 2026
Testing the 2 BIGGEST AI Website Builders

Testing the 2 BIGGEST AI Website Builders

July 17, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!