• bitcoinBitcoin(BTC)$79,554.00-1.75%
  • ethereumEthereum(ETH)$2,452.16-2.50%
  • tetherTether(USDT)$1.000.02%
  • binancecoinBNB(BNB)$721.72-0.44%
  • rippleXRP(XRP)$1.40-3.67%
  • usd-coinUSDC(USDC)$1.000.01%
  • solanaSolana(SOL)$101.85-2.11%
  • tronTRON(TRX)$0.3317690.65%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.040.54%
  • HyperliquidHyperliquid(HYPE)$84.07-3.54%
  • zcashZcash(ZEC)$1,026.848.21%
  • dogecoinDogecoin(DOGE)$0.084590-3.24%
  • RainRain(RAIN)$0.016445-4.09%
  • moneroMonero(XMR)$531.755.38%
  • USDSUSDS(USDS)$1.00-0.02%
  • chainlinkChainlink(LINK)$11.69-2.03%
  • whitebitWhiteBIT Coin(WBT)$73.12-1.08%
  • leo-tokenLEO Token(LEO)$9.22-1.05%
  • cardanoCardano(ADA)$0.210456-6.23%
  • stellarStellar(XLM)$0.179877-2.19%
  • bitcoin-cashBitcoin Cash(BCH)$247.14-3.46%
  • daiDai(DAI)$1.000.01%
  • Ethena USDeEthena USDe(USDE)$1.000.01%
  • CantonCanton(CC)$0.107573-4.09%
  • USD1USD1(USD1)$1.000.01%
  • litecoinLitecoin(LTC)$52.322.31%
  • uniswapUniswap(UNI)$6.27-3.28%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.401.97%
  • hedera-hashgraphHedera(HBAR)$0.0788630.69%
  • Global DollarGlobal Dollar(USDG)$1.00-0.01%
  • avalanche-2Avalanche(AVAX)$7.39-1.74%
  • suiSui(SUI)$0.76-1.73%
  • shiba-inuShiba Inu(SHIB)$0.000005-2.60%
  • nearNEAR Protocol(NEAR)$2.2314.03%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.01%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • tether-goldTether Gold(XAUT)$4,427.79-0.86%
  • crypto-com-chainCronos(CRO)$0.055782-3.71%
  • Circle USYCCircle USYC(USYC)$1.140.04%
  • MemeCoreMemeCore(M)$1.139.23%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • okbOKB(OKB)$109.07-0.03%
  • BittensorBittensor(TAO)$228.42-0.15%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.14-0.42%
  • aaveAave(AAVE)$130.57-2.65%
  • AsterAster(ASTER)$0.731.55%
  • pax-goldPAX Gold(PAXG)$4,435.49-0.88%
  • mantleMantle(MNT)$0.570.48%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.056528-2.51%
  • OndoOndo(ONDO)$0.3679490.86%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Thinking Machines debuts Inkling Small open source AI model nearing performance of predecessor at about 1/4 size

July 31, 2026
in AI & Technology
Reading Time: 6 mins read
A A
Thinking Machines debuts Inkling Small open source AI model nearing performance of predecessor at about 1/4 size
ShareShareShareShareShare

Just two weeks after Thinking Machines released Inkling, its first open source AI language model, the well-funded startup led by former OpenAI chief technology officer Mira Murati today introduced Inkling-Small without sacrificing much of any performance — and in fact, the new model surpasses its larger predecessor on several benchmarks.

YOU MAY ALSO LIKE

How To See What’s Taking Up Space On Your Windows PC

The Tetris Company Wants Nothing To Do With The White House’s New Copycat Game

Inkling Small is a 276-billion-parameter multimodal reasoning model with a permissive Apache 2.0 license that comes within a single point of its larger sibling on the third-party Artificial Analysis Intelligence Index, despite the original Inkling being 975 billion parameters (internal model settings). It accepts text, image and audio inputs, produces text, and supports a context window of up to one million tokens.

Inkling Small uses 12 billion active parameters per token, compared with Inkling’s 41 billion active parameters, while preserving much of the flagship’s coding, reasoning and multimodal performance.

For enterprises, the appeal is not simply that Inkling-Small is smaller. It is that developers appear to give up relatively little capability while reducing the model’s compute requirements, inference costs and deployment footprint.

The model remains far too large for a laptop or conventional workstation, but it is materially easier to operate than the 3.5X larger flagship, making it a good fit for enterprises with some — but not a lot — of their own graphics processing units (GPUs).

Thinking Machines has released the full weights on Hugging Face and added support for fine-tuning through its Tinker model training application programming interface (API).

At launch, the company is advertising a limited-time 50% discount, bringing API pricing for the standard 64K-context Inkling-Small model to $0.58 per million prefill (input) tokens, $1.44 per million sampled (output) tokens, and $1.73 per million training tokens, with cached prefill requests priced at $0.116 per million tokens. A 256K-context variant is also available at higher rates.

Nearly the same performance at a quarter the size

Artificial Analysis assigned Inkling-Small a score of 40 on its Intelligence Index, compared with 41 for Inkling.

That result is notable because Inkling-Small has 276 billion total parameters and 12 billion active parameters, while Inkling has 975 billion total parameters and 41 billion active parameters.

Artificial Analysis also reported that no open-weight model at Inkling-Small’s size or smaller scored higher on the index.

The model does more than merely approach the flagship’s aggregate score. On several evaluations, it surpasses Inkling.

Thinking Machines reports that Inkling-Small scores 80.2% on SWE-bench Verified, compared with Inkling’s 77.6%, and 64.7% on Terminal Bench 2.1, compared with 63.8% for the larger model. It also edges ahead on SciCode, Humanity’s Last Exam, GPQA Diamond and CritPt.

The gains are not universal. Inkling retains a clear advantage on factual knowledge and some agentic tasks. Inkling-Small scores 15.5% on τ³-Banking, compared with 23.7% for Inkling, and its AA Omniscience score is negative, reflecting weaker factual coverage even though its reported hallucination rate is slightly lower.

That tradeoff matters for enterprises. Inkling-Small may be attractive for coding assistants, tool-use systems, retrieval-augmented generation, document analysis and multimodal workflows, but organizations using it for high-stakes factual tasks will still need retrieval, verification and human review.

How a 276B model uses only 12B parameters at a time

Inkling-Small is a sparse Mixture-of-Experts model. According to the model card published by Thinking Machines, its 42-layer decoder routes each token to six of 256 specialized experts, along with two shared experts that remain active for every token.

That architecture helps explain the distinction between the model’s 276 billion total parameters and its 12 billion active parameters. The system retains a large pool of learned capacity but activates only a fraction of it during each inference step.

It is also natively multimodal. Images, audio and text are projected into a shared representation and processed jointly by the decoder rather than being handled through completely separate external systems. Thinking Machines lists coding assistants, agentic applications, chatbots, RAG systems and other multimodal applications among its intended uses.

The company also supports variable reasoning effort, allowing developers to increase or reduce the model’s test-time compute depending on the difficulty of the task. That gives engineering teams a direct way to balance quality, latency and cost across different workloads.

Unfortunately, small does not mean it runs on a laptop

Despite its name, Inkling-Small is not a consumer-scale model.

The standard BF16 checkpoint requires at least 600 GB of aggregate GPU memory, according to Thinking Machines. The company lists two supported configurations: 4x NVIDIA B300 GPUs or 8x NVIDIA H200 GPUs.

A quantized NVFP4 checkpoint lowers the requirement to roughly 180 GB of aggregate VRAM. Thinking Machines says that version can run in W4A4 mode on a single NVIDIA B300, or in W4A16 mode on two H200 GPUs.

That rules out ordinary laptops, MacBooks, desktop gaming PCs and most developer workstations. Even heavily equipped local systems generally fall far short of the required memory.

The practical deployment targets are enterprise GPU servers, cloud clusters and specialized inference providers. The “Small” label is therefore relative to Inkling, not to the broader universe of local models.

Still, the reduction is meaningful. A model that approaches Inkling’s performance while needing substantially less aggregate memory can lower hosting costs, make capacity planning easier and widen the group of organizations capable of self-hosting it.

For companies that want control over data, model behavior and fine-tuning, that smaller footprint may be more important than chasing the highest possible benchmark score.

And of course, it being open source means that it will no doubt be rapidly quantized (made less precise but requiring less compute) and likely blended with other models to be made even smaller for consumer-grade hardware.

Apache 2.0 is the gold standard for enterprise open source models

The licensing may be as important as the benchmarks.

Inkling-Small is released under Apache 2.0, one of the software industry’s most familiar permissive licenses. It generally allows organizations to use, modify, fine-tune, redistribute and commercialize the model, including inside proprietary products, provided they comply with the license’s notice and attribution requirements.

That gives enterprises far more legal flexibility than many custom “open” AI licenses, which may include revenue thresholds, branding obligations, use restrictions or separate conditions for large-scale commercial deployment.

The distinction is increasingly relevant as more AI companies publish model weights without using a conventional open-source license.

Chinese AI darling Moonshot for example, made the weights of its frontier class Kimi K3 model available earlier this week under a custom “open” license that includes additional commercial conditions rather than the comparatively straightforward terms of Apache 2.0.

For legal, procurement and platform teams, that difference can materially simplify adoption. Apache 2.0 does not eliminate the need to review acceptable-use policies, data provenance, regulatory exposure or downstream safety obligations. But it gives organizations a clearer starting point for building internal systems, shipping commercial products and maintaining modified versions of the model.

A more repeatable model-development pipeline

Inkling-Small also shows how quickly Thinking Machines has turned its first large model release into a repeatable engineering process.

Thinking Machines researcher Horace He contrasted the two launches in a post on X:

“Whereas I felt like it took a village to release Inkling, Inkling-Small felt much more routine 😆 We just took the pipeline used for Inkling, passed in a smaller model, and voila — new model! Inkling Small benefited quite a bit vs Inkling from some minor improvements, but there’s still so much more left in the tank…”

The comment suggests the company is no longer treating each model as a one-off research project. Instead, it is building a reusable pipeline for pre-training, post-training, reinforcement learning, evaluation and release.

Thinking Machines says Inkling-Small benefited from an improved pre-training data mix, changes to the machine-learning recipe and on-policy distillation using Inkling as a teacher. The team then continued agentic coding reinforcement learning for two weeks.

Mira Murati emphasized the same point in her own post, describing Inkling-Small as comparable to Inkling at one quarter of the size and highlighting that the weights were open and fine-tunable on Tinker immediately.

How enterprises and AI builders should think about Inkling Small

The company is also distributing full BF16 and NVFP4 checkpoints and supporting deployment through SGLang, vLLM, TokenSpeed, Unsloth and Hugging Face tooling.

That combination gives developers several deployment paths: use an API, fine-tune through Tinker, rely on a third-party inference provider, or operate the model on private infrastructure.

Inkling-Small is not a model that most individuals will download and run locally. But for businesses deciding between a very large flagship and a more manageable open-weight system, it presents a compelling compromise: nearly the same measured intelligence, stronger results on several coding and reasoning tasks, lower token pricing, a smaller hardware footprint and a license that permits broad commercial development.

The broader signal may be just as important. Thinking Machines is showing that Inkling was not a one-time release. The company is already compressing its model family, refining its training pipeline and moving toward a cadence in which open-weight multimodal systems can be produced, improved and deployed more routinely.

Credit: Source link

ShareTweetSendSharePin

Related Posts

How To See What’s Taking Up Space On Your Windows PC
AI & Technology

How To See What’s Taking Up Space On Your Windows PC

September 4, 2026
The Tetris Company Wants Nothing To Do With The White House’s New Copycat Game
AI & Technology

The Tetris Company Wants Nothing To Do With The White House’s New Copycat Game

September 4, 2026
OpenAI Commits B to Frontline Cyber Defense, Launches MS-ISAC Pilot – Unite.AI
AI & Technology

OpenAI Commits $1B to Frontline Cyber Defense, Launches MS-ISAC Pilot – Unite.AI

September 4, 2026
Flock Cameras Are Officially Banned On State Roads In Florida
AI & Technology

Flock Cameras Are Officially Banned On State Roads In Florida

September 4, 2026
Next Post
Why Did $META Crash 10% After Reporting Earnings?

Why Did $META Crash 10% After Reporting Earnings?

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
The AI Boom Isn’t Peaking — Here’s Where Smart Money Is Going

The AI Boom Isn’t Peaking — Here’s Where Smart Money Is Going

September 4, 2026
Wall Street banks tell Big Law to cut fees as AI speeds up legal work

Wall Street banks tell Big Law to cut fees as AI speeds up legal work

September 1, 2026
Bamboo Insurance Services Posts Strong Growth Before IPO Filing (Pending:BMB)

Bamboo Insurance Services Posts Strong Growth Before IPO Filing (Pending:BMB)

September 1, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!