• bitcoinBitcoin(BTC)$84,772.000.83%
  • ethereumEthereum(ETH)$2,716.661.03%
  • tetherTether(USDT)$1.000.00%
  • binancecoinBNB(BNB)$778.360.66%
  • rippleXRP(XRP)$1.54-0.77%
  • usd-coinUSDC(USDC)$1.000.00%
  • solanaSolana(SOL)$124.463.31%
  • tronTRON(TRX)$0.333956-0.96%
  • zcashZcash(ZEC)$1,660.348.34%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.063.68%
  • HyperliquidHyperliquid(HYPE)$93.171.52%
  • dogecoinDogecoin(DOGE)$0.0977660.17%
  • chainlinkChainlink(LINK)$14.382.47%
  • moneroMonero(XMR)$557.430.31%
  • whitebitWhiteBIT Coin(WBT)$84.670.90%
  • USDSUSDS(USDS)$1.000.01%
  • cardanoCardano(ADA)$0.2575831.14%
  • RainRain(RAIN)$0.0127187.07%
  • leo-tokenLEO Token(LEO)$9.010.86%
  • stellarStellar(XLM)$0.2186750.56%
  • nearNEAR Protocol(NEAR)$5.3911.34%
  • bitcoin-cashBitcoin Cash(BCH)$343.902.21%
  • uniswapUniswap(UNI)$10.034.12%
  • litecoinLitecoin(LTC)$71.87-2.27%
  • CantonCanton(CC)$0.136292-0.19%
  • suiSui(SUI)$1.2810.03%
  • Ethena USDeEthena USDe(USDE)$1.000.01%
  • avalanche-2Avalanche(AVAX)$11.114.34%
  • daiDai(DAI)$1.00-0.01%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.609.84%
  • USD1USD1(USD1)$1.000.02%
  • hedera-hashgraphHedera(HBAR)$0.0953521.26%
  • BittensorBittensor(TAO)$334.647.34%
  • shiba-inuShiba Inu(SHIB)$0.0000061.37%
  • crypto-com-chainCronos(CRO)$0.0676433.37%
  • Global DollarGlobal Dollar(USDG)$1.00-0.01%
  • BitwayBitway(BTW)$1.0718.42%
  • MemeCoreMemeCore(M)$1.23-0.07%
  • paypal-usdPayPal USD(PYUSD)$1.000.01%
  • EthenaEthena(ENA)$0.271498-1.52%
  • tether-goldTether Gold(XAUT)$4,279.32-0.04%
  • OndoOndo(ONDO)$0.54-0.67%
  • okbOKB(OKB)$121.640.14%
  • quant-networkQuant(QNT)$174.0867.24%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • aaveAave(AAVE)$156.350.97%
  • Circle USYCCircle USYC(USYC)$1.140.00%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.15-0.10%
  • mantleMantle(MNT)$0.68-4.29%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

NVIDIA AI Released Jet-Nemotron: 53x Faster Hybrid-Architecture Language Model Series that Translates to a 98% Cost Reduction for Inference at Scale

August 27, 2025
in AI & Technology
Reading Time: 8 mins read
A A
NVIDIA AI Released Jet-Nemotron: 53x Faster Hybrid-Architecture Language Model Series that Translates to a 98% Cost Reduction for Inference at Scale
ShareShareShareShareShare

NVIDIA researchers have shattered the longstanding efficiency hurdle in large language model (LLM) inference, releasing Jet-Nemotron—a family of models (2B and 4B) that delivers up to 53.6× higher generation throughput than leading full-attention LLMs while matching, or even surpassing, their accuracy. Most importantly, this breakthrough isn’t the result of a new pre-training run from scratch, but rather a retrofit of existing, pre-trained models using a novel technique called Post Neural Architecture Search (PostNAS). The implications are transformative for businesses, practitioners, and researchers alike.

The Need for Speed in Modern LLMs

While today’s state-of-the-art (SOTA) LLMs, like Qwen3, Llama3.2, and Gemma3, have set new benchmarks for accuracy and flexibility, their O(n²) self-attention mechanism incurs exorbitant costs—both in compute and memory—especially for long-context tasks. This makes them expensive to deploy at scale and nearly impossible to run on edge or memory-constrained devices. Efforts to replace full-attention Transformers with more efficient architectures (Mamba2, GLA, RWKV, etc.) have struggled to close the accuracy gap, until now.

YOU MAY ALSO LIKE

AI Coding Agents for Enterprise: IP Indemnity, Data Residency and 500-Seat Cost Compared

Why We Won’t Know How Visible The iPhone Duo’s Crease Is For A Long Time

https://arxiv.org/abs/2508.15884v1?

PostNAS: A Surgical, Capital-Efficient Overhaul

The core innovation is PostNAS: a neural architecture search pipeline designed specifically for efficiently retrofitting pre-trained models. Here’s how it works:

  • Freeze the Knowledge: Start with a SOTA full-attention model (like Qwen2.5). Freeze its MLP layers—this preserves the model’s learned intelligence and greatly reduces training cost.
  • Surgical Replacement: Replace computationally expensive full-attention (Transformers) with JetBlock, a new, hardware-efficient linear attention block designed for NVIDIA’s latest GPUs.
  • Hybrid, Hardware-Aware Design: Use super-network training and beam search to automatically determine the optimal placement and minimal set of full-attention layers necessary to preserve accuracy on key tasks (retrieval, math, MMLU, coding, etc.). This step is task-specific and hardware-aware: the search maximizes throughput for target hardware, not just parameter count.
  • Scale and Deploy: The result is a hybrid-architecture LLM that inherits the backbone intelligence of the original model but slashes latency and memory footprint.

JetBlock is particularly noteworthy: it introduces dynamic causal convolution kernels conditioned on input (unlike static kernels in prior linear attention blocks) and removes redundant convolutions for streamlined efficiency. With hardware-aware hyperparameter search, it not only keeps pace with prior linear attention designs in throughput, but actually boosts accuracy.

https://arxiv.org/abs/2508.15884v1?

Jet-Nemotron: Performance by the Numbers

The key metrics from NVIDIA’s technical paper are staggering:

Model MMLU-Pro Acc. Generation Throughput (tokens/s, H100) KV Cache Size (MB, 64K context) Notes
Qwen3-1.7B-Base 37.8 61 7,168 Full-attention baseline
Jet-Nemotron-2B 39.0 2,885 154 47× throughput, 47× smaller cache
Jet-Nemotron-4B 44.2 1,271 258 21× throughput, still SOTA acc.
Mamba2-2.7B 8.6 2,507 80 All-linear, much lower accuracy
RWKV7-1.5B 13.4 3,050 24 All-linear, much lower accuracy
DeepSeek-V3-Small (MoE) — — — 2.2B activated, 15B total, lower acc.

Jet-Nemotron-2B matches or exceeds Qwen3-1.7B-Base on every major benchmark—math, commonsense, coding, retrieval, long-context—while delivering 47× higher generation throughput.

This isn’t a small gain: a 53.6× speedup in decoding at 256K context length means a 98% reduction in inference cost for the same volume of tokens. Prefilling speedups are also dramatic: 6.14× faster at 256K context.

Memory footprint shrinks by 47× (154MB cache vs. 7,168MB for Qwen3-1.7B-Base). This is a game-changer for edge deployment: Jet-Nemotron-2B is 8.84× and 6.5× faster than Qwen2.5-1.5B on Jetson Orin and RTX 3090, respectively.

https://arxiv.org/abs/2508.15884v1?

Applications

For Business Leaders: Better ROI $$

  • Inference at scale is now affordable. A 53× throughput gain means dollar-for-dollar, you can serve 53× more users—or slash hosting costs by 98%.
  • Operational efficiency is transformed: latency drops, batch sizes grow, and memory constraints vanish. Cloud providers can offer SOTA AI at commodity prices.
  • The AI business model reshapes: Tasks once too expensive (real-time document AI, long-context agents, on-device copilots) suddenly become viable.

For Practitioners: SOTA on the Edge

  • Forget about quantization, distillation, or pruning compromises. Jet-Nemotron’s tiny KV cache (154MB) and 2B parameters fit on Jetson Orin, RTX 3090, and even mobile chips—no more offloading to the cloud.
  • No retraining, no data pipeline changes: Just retrofitting. Your existing Qwen, Llama, or Gemma checkpoints can be upgraded without losing accuracy.
  • Real-world AI services (search, copilots, summarization, coding) are now instant and scalable.

For Researchers: Lower Barrier, Higher Innovation

  • PostNAS slashes the cost of LLM architecture innovation. Instead of months and millions on pre-training, architecture search happens on frozen backbone models in a fraction of the time.
  • Hardware-aware NAS is the future: The Jet-Nemotron process considers KV cache size (not just parameters) as the critical factor for real-world speed. This is a paradigm shift in how we measure and optimize efficiency.
  • The community can iterate faster: PostNAS is a rapid testbed. If a new attention block works here, it’s worth pre-training; if not, it’s filtered out before the big spend.

Summary

The open-sourcing of Jet-Nemotron and JetBlock (code on GitHub) means the broader AI ecosystem can now retrofit their models for unprecedented efficiency. PostNAS is not a one-off trick: it’s a general-purpose framework for accelerating any Transformer, lowering the cost of future breakthroughs.


Check out the Paper and GitHub Page. Feel free to check out our GitHub Page for Tutorials, Codes and Notebooks. Also, feel free to follow us on Twitter and don’t forget to join our 100k+ ML SubReddit and Subscribe to our Newsletter.

The post NVIDIA AI Released Jet-Nemotron: 53x Faster Hybrid-Architecture Language Model Series that Translates to a 98% Cost Reduction for Inference at Scale appeared first on MarkTechPost.

Credit: Source link

ShareTweetSendSharePin

Related Posts

AI Coding Agents for Enterprise: IP Indemnity, Data Residency and 500-Seat Cost Compared
AI & Technology

AI Coding Agents for Enterprise: IP Indemnity, Data Residency and 500-Seat Cost Compared

September 27, 2026
Why We Won’t Know How Visible The iPhone Duo’s Crease Is For A Long Time
AI & Technology

Why We Won’t Know How Visible The iPhone Duo’s Crease Is For A Long Time

September 27, 2026
How Powerful Of A Power Bank Do You Need To Safely Charge A Laptop?
AI & Technology

How Powerful Of A Power Bank Do You Need To Safely Charge A Laptop?

September 27, 2026
Your Old GPU Could Be Worth More Than You Think
AI & Technology

Your Old GPU Could Be Worth More Than You Think

September 26, 2026
Next Post
Climbing games are so hot right now

Climbing games are so hot right now

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
New hearing in Idaho college murders case

New hearing in Idaho college murders case

September 22, 2026
U.S. imposes 50% tariffs on some Canadian goods

U.S. imposes 50% tariffs on some Canadian goods

September 25, 2026
Florida lifeguard saves his own family after boat crash

Florida lifeguard saves his own family after boat crash

September 24, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!