• bitcoinBitcoin(BTC)$77,817.00-3.46%
  • ethereumEthereum(ETH)$2,444.63-2.74%
  • tetherTether(USDT)$1.000.00%
  • binancecoinBNB(BNB)$691.31-3.23%
  • rippleXRP(XRP)$1.39-4.51%
  • usd-coinUSDC(USDC)$1.000.00%
  • solanaSolana(SOL)$104.26-3.86%
  • tronTRON(TRX)$0.3410740.52%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.042.09%
  • HyperliquidHyperliquid(HYPE)$80.39-4.37%
  • zcashZcash(ZEC)$806.880.10%
  • dogecoinDogecoin(DOGE)$0.085329-4.31%
  • RainRain(RAIN)$0.0176792.18%
  • USDSUSDS(USDS)$1.000.01%
  • leo-tokenLEO Token(LEO)$9.642.03%
  • moneroMonero(XMR)$470.971.26%
  • chainlinkChainlink(LINK)$11.42-4.16%
  • whitebitWhiteBIT Coin(WBT)$71.76-3.31%
  • cardanoCardano(ADA)$0.202651-5.72%
  • stellarStellar(XLM)$0.178793-4.42%
  • bitcoin-cashBitcoin Cash(BCH)$247.47-8.03%
  • daiDai(DAI)$1.000.00%
  • CantonCanton(CC)$0.110388-2.98%
  • USD1USD1(USD1)$1.00-0.01%
  • Ethena USDeEthena USDe(USDE)$1.000.00%
  • litecoinLitecoin(LTC)$49.15-1.91%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.36-3.68%
  • Global DollarGlobal Dollar(USDG)$1.00-0.01%
  • hedera-hashgraphHedera(HBAR)$0.075545-4.15%
  • avalanche-2Avalanche(AVAX)$7.30-2.98%
  • shiba-inuShiba Inu(SHIB)$0.000005-5.23%
  • suiSui(SUI)$0.74-5.61%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.01%
  • Circle USYCCircle USYC(USYC)$1.140.03%
  • uniswapUniswap(UNI)$4.43-6.53%
  • crypto-com-chainCronos(CRO)$0.056782-6.96%
  • tether-goldTether Gold(XAUT)$4,457.92-2.49%
  • nearNEAR Protocol(NEAR)$1.81-6.57%
  • MemeCoreMemeCore(M)$1.04-6.92%
  • Ripple USDRipple USD(RLUSD)$1.00-0.03%
  • okbOKB(OKB)$109.71-3.86%
  • BittensorBittensor(TAO)$236.74-5.52%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.14-0.10%
  • pax-goldPAX Gold(PAXG)$4,461.38-2.47%
  • aaveAave(AAVE)$121.81-5.86%
  • AsterAster(ASTER)$0.69-2.88%
  • Pump.funPump.fun(PUMP)$0.004625-3.69%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.057666-5.51%
  • OndoOndo(ONDO)$0.353635-7.16%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Alibaba Introduces Qwen3-Max-Thinking, a Test Time Scaled Reasoning Model with Native Tool Use Powering Agentic Workloads

January 29, 2026
in AI & Technology
Reading Time: 5 mins read
A A
Alibaba Introduces Qwen3-Max-Thinking, a Test Time Scaled Reasoning Model with Native Tool Use Powering Agentic Workloads
ShareShareShareShareShare

Qwen3-Max-Thinking is Alibaba’s new flagship reasoning model. It does not only scale parameters, it also changes how inference is done, with explicit control over thinking depth and built in tools for search, memory, and code execution.

https://qwen.ai/blog?id=qwen3-max-thinking

Model scale, data, and deployment

Qwen3-Max-Thinking is a trillion-parameter MoE flagship LLM pretrained on 36T tokens and built on the Qwen3 family as the top tier reasoning model. The model targets long horizon reasoning and code, not only casual chat. It runs with a context window of 260k tokens, which supports repository scale code, long technical reports, and multi document analysis within a single prompt.

YOU MAY ALSO LIKE

How To Use Your Old Android Phone As A Wi-Fi Extender

When Content Is Infinite, Point of View Becomes the Scarce Asset – Unite.AI

Qwen3-Max-Thinking is a closed model served through Qwen-Chat and Alibaba Cloud Model Studio with an OpenAI compatible HTTP API. The same endpoint can be called in a Claude style tool schema, so existing Anthropic or Claude Code flows can swap in Qwen3-Max-Thinking with minimal changes. There are no public weights, so usage is API based, which matches its positionin

Smart Test Time Scaling and experience cumulative reasoning

Most large language models improve reasoning by simple test time scaling, for example best of N sampling with several parallel chains of thought. That approach increases quality but cost grows almost linearly with the number of samples. Qwen3-Max-Thinking introduces an experience cumulative, multi round test time scaling strategy.

Instead of only sampling more in parallel, the model iterates within a single conversation, reusing intermediate reasoning traces as structured experience. After each round, it extracts useful partial conclusions, then focuses subsequent computation on unresolved parts of the question. This process is controlled by an explicit thinking budget that developers can adjust via API parameters such as enable_thinking and additional configuration fields.

The reported effect is that accuracy rises without a proportional increase in token count. For example, Qwen’s own ablations show GPQA Diamond increasing from around 90 level accuracy to about 92.8, and LiveCodeBench v6 rising from about 88.0 to 91.4 under the experience cumulative strategy at similar token budgets. This is important because it means higher reasoning quality can be driven by more efficient scheduling of compute, not only by more samples.

Native agent stack with Adaptive Tool Use

Qwen3-Max-Thinking integrates three tools as first class capabilities: Search, Memory, and a Code Interpreter. Search connects to web retrieval so the model can fetch fresh pages, extract content, and ground its answers. Memory stores user or session specific state, which supports personalized reasoning over longer workflows. The Code Interpreter executes Python, which allows numeric verification, data transforms, and program synthesis with runtime checks.

The model uses Adaptive Tool Use to decide when to invoke these tools during a conversation. Tool calls are interleaved with internal thinking segments, rather than being orchestrated by an external agent. This design reduces the need for separate routers or planners and tends to reduce hallucinations, because the model can explicitly fetch missing information or verify calculations instead of guessing.

Tool ability is also benchmarked. On Tau² Bench, which measures function calling and tool orchestration, Qwen3-Max-Thinking reports a score of 82.1, comparable with other frontier models in this category.

Benchmark profile across knowledge, reasoning, and search

On 19 public benchmarks, Qwen3-Max-Thinking is positioned at or near the same level as GPT 5.2 Thinking, Claude Opus 4.5, and Gemini 3 Pro. For knowledge tasks, reported scores include 85.7 on MMLU-Pro, 92.8 on MMLU-Redux, and 93.7 on C-Eval, where Qwen leads the group on Chinese language evaluation.

For hard reasoning, it records 87.4 on GPQA, 98.0 on HMMT Feb 25, 94.7 on HMMT Nov 25, and 83.9 on IMOAnswerBench, which puts it in the top tier of current math and science models. On coding and software engineering it reaches 85.9 on LiveCodeBench v6 and 75.3 on SWE Verified.

In the base HLE configuration Qwen3-Max-Thinking scores 30.2, below Gemini 3 Pro at 37.5 and GPT 5.2 Thinking at 35.5. In a tool enabled HLE setup, the official comparison table that includes web search integration shows Qwen3-Max-Thinking at 49.8, ahead of GPT 5.2 Thinking at 45.5 and Gemini 3 Pro at 45.8. With its most aggressive experience cumulative test time scaling configuration on HLE with tools, Qwen3-Max-Thinking reaches 58.3 while GPT 5.2 Thinking remains at 45.5, although that higher number is for a heavier inference mode than the standard comparison table.

Key Takeaways

  • Qwen3-Max-Thinking is a closed, API only flagship reasoning model from Alibaba, built on a more than 1 trillion parameter backbone trained on about 36 trillion tokens with a 262144 token context window.
  • The model introduces experience cumulative test time scaling, where it reuses intermediate reasoning across multiple rounds, improving benchmarks such as GPQA Diamond and LiveCodeBench v6 at similar token budgets.
  • Qwen3-Max-Thinking integrates Search, Memory, and a Code Interpreter as native tools and uses Adaptive Tool Use so the model itself decides when to browse, recall state, or execute Python during a conversation.
  • On public benchmarks it reports competitive scores with GPT 5.2 Thinking, Claude Opus 4.5, and Gemini 3 Pro, including strong results on MMLU Pro, GPQA, HMMT, IMOAnswerBench, LiveCodeBench v6, SWE Bench Verified, and Tau² Bench..

Check out the API and Technical details. Also, feel free to follow us on Twitter and don’t forget to join our 100k+ ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well.

The post Alibaba Introduces Qwen3-Max-Thinking, a Test Time Scaled Reasoning Model with Native Tool Use Powering Agentic Workloads appeared first on MarkTechPost.

Credit: Source link

ShareTweetSendSharePin

Related Posts

How To Use Your Old Android Phone As A Wi-Fi Extender
AI & Technology

How To Use Your Old Android Phone As A Wi-Fi Extender

August 28, 2026
When Content Is Infinite, Point of View Becomes the Scarce Asset – Unite.AI
AI & Technology

When Content Is Infinite, Point of View Becomes the Scarce Asset – Unite.AI

August 28, 2026
Anthropic Reports Claude Agents Mitigated Ten Alignment Failures – Unite.AI
AI & Technology

Anthropic Reports Claude Agents Mitigated Ten Alignment Failures – Unite.AI

August 28, 2026
Early Leak Of NVIDIA’s DLSS 5 Has An Uncanny Valley Problem
AI & Technology

Early Leak Of NVIDIA’s DLSS 5 Has An Uncanny Valley Problem

August 28, 2026
Next Post
Machado presents her Nobel Peace Prize to Trump

Machado presents her Nobel Peace Prize to Trump

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Who Killed Tupac Shakur? And Can Democratic Socialists Win in November? – Aug. 10

Who Killed Tupac Shakur? And Can Democratic Socialists Win in November? – Aug. 10

August 26, 2026
If Waymo Cars Are Level 4 Automation, What Does It Take To Be A Level 5?

If Waymo Cars Are Level 4 Automation, What Does It Take To Be A Level 5?

August 22, 2026
Jinny Lu is crowned the World’s Ugliest Dog

Jinny Lu is crowned the World’s Ugliest Dog

August 22, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!