• bitcoinBitcoin(BTC)$78,477.00-0.79%
  • ethereumEthereum(ETH)$2,472.320.30%
  • tetherTether(USDT)$1.000.01%
  • binancecoinBNB(BNB)$698.21-0.13%
  • rippleXRP(XRP)$1.38-6.11%
  • usd-coinUSDC(USDC)$1.000.01%
  • solanaSolana(SOL)$96.76-1.66%
  • tronTRON(TRX)$0.335749-1.22%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.01-2.33%
  • HyperliquidHyperliquid(HYPE)$81.03-1.46%
  • dogecoinDogecoin(DOGE)$0.085097-4.04%
  • zcashZcash(ZEC)$783.41-1.44%
  • RainRain(RAIN)$0.017633-2.27%
  • USDSUSDS(USDS)$1.00-0.03%
  • whitebitWhiteBIT Coin(WBT)$72.57-0.61%
  • leo-tokenLEO Token(LEO)$9.23-1.35%
  • chainlinkChainlink(LINK)$11.29-2.23%
  • moneroMonero(XMR)$430.73-2.72%
  • cardanoCardano(ADA)$0.206225-4.09%
  • stellarStellar(XLM)$0.180120-5.35%
  • bitcoin-cashBitcoin Cash(BCH)$261.85-2.36%
  • CantonCanton(CC)$0.117447-3.33%
  • daiDai(DAI)$1.000.00%
  • USD1USD1(USD1)$1.000.01%
  • Ethena USDeEthena USDe(USDE)$1.000.00%
  • litecoinLitecoin(LTC)$49.92-2.24%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.39-5.89%
  • hedera-hashgraphHedera(HBAR)$0.077547-2.92%
  • Global DollarGlobal Dollar(USDG)$1.000.01%
  • avalanche-2Avalanche(AVAX)$7.25-3.20%
  • shiba-inuShiba Inu(SHIB)$0.000005-3.74%
  • suiSui(SUI)$0.74-5.35%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • crypto-com-chainCronos(CRO)$0.058318-2.07%
  • tether-goldTether Gold(XAUT)$4,583.82-1.17%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.02%
  • uniswapUniswap(UNI)$4.25-1.74%
  • MemeCoreMemeCore(M)$1.12-5.65%
  • nearNEAR Protocol(NEAR)$1.83-2.82%
  • okbOKB(OKB)$110.92-3.15%
  • BittensorBittensor(TAO)$228.65-2.84%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.14-0.01%
  • Ripple USDRipple USD(RLUSD)$1.000.01%
  • pax-goldPAX Gold(PAXG)$4,589.92-1.22%
  • aaveAave(AAVE)$123.73-3.49%
  • AsterAster(ASTER)$0.70-2.01%
  • Pump.funPump.fun(PUMP)$0.0048136.21%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.058031-1.44%
  • OndoOndo(ONDO)$0.361997-2.73%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

MBZUAI Releases K2 Think V2: A Fully Sovereign 70B Reasoning Model For Math, Code, And Science

January 28, 2026
in AI & Technology
Reading Time: 6 mins read
A A
MBZUAI Releases K2 Think V2: A Fully Sovereign 70B Reasoning Model For Math, Code, And Science
ShareShareShareShareShare

Can a fully sovereign open reasoning model match state of the art systems when every part of its training pipeline is transparent. Researchers from Mohamed bin Zayed University of Artificial Intelligence (MBZUAI) release K2 Think V2, a fully sovereign reasoning model designed to test how far open and fully documented pipelines can push long horizon reasoning on math, code, and science when the entire stack is open and reproducible. K2 Think V2 takes the 70 billion parameter K2 V2 Instruct base model and applies a carefully engineered reinforcement learning approach to turn it into a high precision reasoning model that remains fully open in both weights and data.

https://arxiv.org/pdf/2512.06201

From K2 V2 base model to reasoning specialist

K2 V2 is a dense decoder only transformer with 80 layers, hidden size 8192, and 64 attention heads with grouped query attention and rotary position embeddings. It is trained on around 12 trillion tokens drawn from the TxT360 corpus and related curated datasets that cover web text, math, code, multilingual data, and scientific literature.

YOU MAY ALSO LIKE

Devolver Digital Squares Up To Rockstar With A Mascot Platformer Out The Same Day As GTA 6

Waystar Puts Agentic AI to Work on Claims, Denials, and Patient Bills – Unite.AI

Training proceeds in three phases. Pretraining runs at context length 8192 tokens on natural data to establish robust general knowledge. Mid training then extends context up to 512k tokens using TxT360 Midas, which mixes long documents, synthetic thinking traces, and diverse reasoning behaviors while carefully keeping at least 30 percent short context data in every stage. Finally, supervised fine tuning, called TxT360 3efforts, injects instruction following and structured reasoning signals.

The important point is that K2 V2 is not a generic base model. It is explicitly optimized for long context consistency and exposure to reasoning behaviors during mid training. That makes it a natural foundation for a post training stage that focuses only on reasoning quality, which is exactly what K2 Think V2 does.

Fully sovereign RLVR on GURU dataset

K2 Think V2 is trained with a GRPO style RLVR recipe on top of K2 V2 Instruct. The team uses the Guru dataset, version 1.5, which focuses on math, code, and STEM questions. Guru is derived from permissively licensed sources, expanded in STEM coverage, and decontaminated against key evaluation benchmarks before use. This is important for a sovereign claim, because both the base model data and the RL data are curated and documented by the same institute.

The GRPO setup removes the usual KL and entropy auxiliary losses and uses asymmetric clipping of the policy ratio with the high clip set to 0.28. Training runs fully on policy with temperature 1.2 to increase rollout diversity, global batch size 256, and no micro batching. This avoids off policy corrections that are known to introduce instability in GRPO like training.

RLVR itself runs in two stages. In the first stage, response length is capped at 32k tokens and the model trains for about 200 steps. In the second stage, the maximum response length is increased to 64k tokens and training continues for about 50 steps with the same hyperparameters. This schedule specifically exploits the long context capability inherited from K2 V2 so that the model can practice full chain of thought trajectories rather than short solutions.

https://mbzuai.ac.ae/news/k2-think-v2-a-fully-sovereign-reasoning-model/

Benchmark profile

K2 Think V2 targets reasoning benchmarks rather than purely knowledge benchmarks. On AIME 2025 it reaches pass at 1 of 90.42. On HMMT 2025 it scores 84.79. On GPQA Diamond, a difficult graduate level science benchmark, it reaches 72.98. On SciCode it records 33.00, and on Humanity’s Last Exam it reaches 9.5 under the benchmark settings.

These scores are reported as averages over 16 runs and are directly comparable only within the same evaluation protocol. The MBZUAI team also highlights improvements on IFBench and on the Artificial Analysis evaluation suite, with particular gains in hallucination rate and long context reasoning compared with the previous K2 Think release.

Safety and openness

The research team reports a Safety 4 style analysis that aggregates four safety surfaces. Content and public safety, truthfulness and reliability, and societal alignment all reach macro average risk levels in the low range. Data and infrastructure risks remain higher and are marked as critical, which reflects concerns about sensitive personal information handling rather than model behavior alone. The team states that K2 Think V2 still shares the generic limitations of large language models despite these mitigations. On Artificial Analysis’s Openness Index, K2 Think V2 sits at the frontier together with K2 V2 and Olmo-3.

Key Takeaways

  • K2 Think V2 is a fully sovereign 70B reasoning model: Built on K2 V2 Instruct, with open weights, open data recipes, detailed training logs, and full RL pipeline released via Reasoning360.
  • Base model is optimized for long context and reasoning before RL: K2 V2 is a dense decoder transformer trained on around 12T tokens, with mid training extending context length to 512K tokens and supervised ‘3 efforts’ SFT targeting structured reasoning.
  • Reasoning is aligned using GRPO based RLVR on the Guru dataset: Training uses a 2 stage on policy GRPO setup on Guru v1.5, with asymmetric clipping, temperature 1.2, and response caps at 32K then 64K tokens to learn long chain of thought solutions.
  • Competitive results on hard reasoning benchmarks: K2 Think V2 reports strong pass at 1 scores such as 90.42 on AIME 2025, 84.79 on HMMT 2025, and 72.98 on GPQA Diamond, positioning it as a high precision open reasoning model for math, code, and science.

Check out the Paper, Model Weight, Repo and Technical details. Also, feel free to follow us on Twitter and don’t forget to join our 100k+ ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well.

The post MBZUAI Releases K2 Think V2: A Fully Sovereign 70B Reasoning Model For Math, Code, And Science appeared first on MarkTechPost.

Credit: Source link

ShareTweetSendSharePin

Related Posts

Devolver Digital Squares Up To Rockstar With A Mascot Platformer Out The Same Day As GTA 6
AI & Technology

Devolver Digital Squares Up To Rockstar With A Mascot Platformer Out The Same Day As GTA 6

August 26, 2026
Waystar Puts Agentic AI to Work on Claims, Denials, and Patient Bills – Unite.AI
AI & Technology

Waystar Puts Agentic AI to Work on Claims, Denials, and Patient Bills – Unite.AI

August 26, 2026
Orchestration is the new challenge for CX in the age of AI agents
AI & Technology

Orchestration is the new challenge for CX in the age of AI agents

August 26, 2026
What Would Have to Be True for Agentic Coding to Replace Junior Engineers
AI & Technology

What Would Have to Be True for Agentic Coding to Replace Junior Engineers

August 26, 2026
Next Post
How Denmark is reacting to Trump’s push for Greenland

How Denmark is reacting to Trump's push for Greenland

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Walmart shares sink 8% on slowest sales growth in over 6 years

Walmart shares sink 8% on slowest sales growth in over 6 years

August 20, 2026
The Hardest Promises to Break: Peer Accountability (With Angela Duckworth)

The Hardest Promises to Break: Peer Accountability (With Angela Duckworth)

August 24, 2026
One killed in Indiana home explosion during major storm

One killed in Indiana home explosion during major storm

August 25, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!