• bitcoinBitcoin(BTC)$83,869.00-3.77%
  • ethereumEthereum(ETH)$2,675.68-3.80%
  • tetherTether(USDT)$1.000.00%
  • binancecoinBNB(BNB)$768.12-3.60%
  • rippleXRP(XRP)$1.49-9.43%
  • usd-coinUSDC(USDC)$1.000.00%
  • solanaSolana(SOL)$114.86-3.83%
  • tronTRON(TRX)$0.343671-0.15%
  • zcashZcash(ZEC)$1,509.83-5.67%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.040.37%
  • HyperliquidHyperliquid(HYPE)$91.88-5.86%
  • dogecoinDogecoin(DOGE)$0.093476-9.82%
  • moneroMonero(XMR)$551.36-3.68%
  • whitebitWhiteBIT Coin(WBT)$84.32-3.70%
  • USDSUSDS(USDS)$1.000.01%
  • chainlinkChainlink(LINK)$12.32-6.18%
  • cardanoCardano(ADA)$0.239905-7.78%
  • RainRain(RAIN)$0.012190-7.33%
  • leo-tokenLEO Token(LEO)$8.990.41%
  • stellarStellar(XLM)$0.201527-9.57%
  • bitcoin-cashBitcoin Cash(BCH)$340.850.13%
  • nearNEAR Protocol(NEAR)$4.421.99%
  • uniswapUniswap(UNI)$9.24-11.86%
  • litecoinLitecoin(LTC)$67.916.43%
  • Ethena USDeEthena USDe(USDE)$1.00-0.01%
  • daiDai(DAI)$1.000.01%
  • avalanche-2Avalanche(AVAX)$10.31-7.74%
  • USD1USD1(USD1)$1.000.00%
  • CantonCanton(CC)$0.108488-6.59%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.42-3.74%
  • hedera-hashgraphHedera(HBAR)$0.090483-10.13%
  • suiSui(SUI)$0.96-6.51%
  • shiba-inuShiba Inu(SHIB)$0.000006-7.96%
  • Global DollarGlobal Dollar(USDG)$1.00-0.01%
  • BittensorBittensor(TAO)$286.41-8.91%
  • crypto-com-chainCronos(CRO)$0.061808-9.32%
  • BitwayBitway(BTW)$1.0922.88%
  • MemeCoreMemeCore(M)$1.26-4.22%
  • paypal-usdPayPal USD(PYUSD)$1.000.00%
  • tether-goldTether Gold(XAUT)$4,285.75-1.37%
  • okbOKB(OKB)$119.16-4.73%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • Ripple USDRipple USD(RLUSD)$1.00-0.01%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.14-0.21%
  • mantleMantle(MNT)$0.66-4.36%
  • aaveAave(AAVE)$138.48-7.99%
  • EthenaEthena(ENA)$0.206341-4.64%
  • OndoOndo(ONDO)$0.424119-4.04%
  • polkadotPolkadot(DOT)$1.13-5.97%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Nvidia’s Vera Rubin is months away — Blackwell is getting faster right now

January 9, 2026
in AI & Technology
Reading Time: 4 mins read
A A
Nvidia’s Vera Rubin is months away — Blackwell is getting faster right now
ShareShareShareShareShare

The big news this week from Nvidia, splashed in headlines across all forms of media, was the company’s announcement about its Vera Rubin GPU.

YOU MAY ALSO LIKE

A Coding Guide to TypeSafe AI Jev: Typed Decisions, Calibrated Confidence, and Speculative Fan-Out with a System One Model

Everything Announced At Meta Connect 2026

This week, Nvidia CEO Jensen Huang used his CES keynote to highlight performance metrics for the new chip. According to Huang, the Rubin GPU is capable of 50 PFLOPs of NVFP4 inference and 35 PFLOPs of NVFP4 training performance, representing 5x and 3.5x the performance of Blackwell.

But it won’t be available until the second half of 2026. So what should enterprises be doing now?

Blackwell keeps on getting better

The current, shipping Nvidia GPU architecture is Blackwell, which was announced in 2024 as the successor to Hopper.  Alongside that release, Nvidia emphasized that that its product engineering path also included squeezing as much performance as possible out of the prior Grace Hopper architecture.

It’s a direction that will hold true for Blackwell as well, with Vera Rubin coming later this year.

“We continue to optimize our inference and training stacks for the Blackwell architecture,” Dave Salvator, director of accelerated computing products at Nvidia, told VentureBeat.

In the same week that Vera Rubin was being touted by Nvidia’s CEO as its most powerful GPU ever, the company published new research showing improved Blackwell performance.

How Blackwell performance has improved inference by 2.8x 

Nvidia has been able to increase Blackwell GPU performance by up to 2.8x per GPU in a period of just three short months.

The performance gains come from a series of innovations that have been added to the Nvidia TensorRT-LLM inference engine. These optimizations apply to existing hardware, allowing current Blackwell deployments to achieve higher throughput without hardware changes.

The performance gains are measured on DeepSeek-R1, a 671-billion parameter mixture-of-experts (MoE) model that activates 37 billion parameters per token.

Among the technical innovations that provide the performance boost:

  • Programmatic dependent launch (PDL): Expanded implementation reduces kernel launch latencies, increasing throughput.

  • All-to-all communication: New implementation of communication primitives eliminates an intermediate buffer, reducing memory overhead.

  • Multi-token prediction (MTP): Generates multiple tokens per forward pass rather than one at a time, increasing throughput across various sequence lengths.

  • NVFP4 format: A 4-bit floating point format with hardware acceleration in Blackwell that reduces memory bandwidth requirements while preserving model accuracy.

The optimizations reduce cost per million tokens and allow existing infrastructure to serve higher request volumes at lower latency. Cloud providers and enterprises can scale their AI services without immediate hardware upgrades.

Blackwell has also made training performance gains 

Blackwell is also widely used as a foundational hardware component for training the largest of large language models.

In that respect, Nvidia has also reported significant gains for Blackwell when used for AI training. 

Since its initial launch, the GB200 NVL72 system delivered up to 1.4x higher training performance on the same hardware — a 40% boost achieved in just five months without any hardware upgrades.

The training boost came from a series of updates including:

  • Optimized training recipes. Nvidia engineers developed sophisticated training recipes that effectively leverage NVFP4 precision. Initial Blackwell submissions used FP8 precision, but the transition to NVFP4-optimized recipes unlocked substantial additional performance from the existing silicon.

  • Algorithmic refinements. Continuous software stack enhancements and algorithmic improvements enabled the platform to extract more performance from the same hardware, demonstrating ongoing innovation beyond initial deployment.

Double-down on Blackwell or wait for Vera Rubin?

Salvator noted that the high-end Blackwell Ultra is a market-leading platform purpose-built to run state-of-the-art AI models and applications. 

He added that the Nvidia Rubin platform will extend the company’s market leadership and enable the next generation of MoEs to power a new class of applications to take AI innovation even further.

Salvator explained that the Vera Rubin is built to address the growing demand in compute created by the continuing growth in model size and reasoning token generation from leading models such as MoE.  

 “Blackwell and Rubin can serve the same models, but the difference is the performance, efficiency and token cost,” he said.

According to Nvidia’s early testing results, compared to Blackwell, Rubin can train large MoE models in a quarter the number of GPUs, inference token generation with 10X more throughput per watt, and inference at 1/10th the cost per token.

“Better token throughput performance and efficiency, means newer models can be built with more reasoning capability and faster agent-to-agent interaction, creating better intelligence at lower cost,” Salvator said.

What it all means for enterprise AI builders

For enterprises deploying AI infrastructure today, current investments in Blackwell remain sound despite Vera Rubin’s arrival later this year.

Organizations with existing Blackwell deployments can immediately capture the 2.8x inference improvement and 1.4x training boost by updating to the latest TensorRT-LLM versions — delivering real cost savings without capital expenditure. For those planning new deployments in the first half of 2026, proceeding with Blackwell makes sense. Waiting six months means delaying AI initiatives and potentially falling behind competitors already deploying today.

However, enterprises planning large-scale infrastructure buildouts for late 2026 and beyond should factor Vera Rubin into their roadmaps. The 10x improvement in throughput per watt and 1/10th cost per token represent transformational economics for AI operations at scale.

The smart approach is phased deployment: Leverage Blackwell for immediate needs while architecting systems that can incorporate Vera Rubin when available. Nvidia’s continuous optimization model means this isn’t a binary choice; enterprises can maximize value from current deployments without sacrificing long-term competitiveness.

Credit: Source link

ShareTweetSendSharePin

Related Posts

A Coding Guide to TypeSafe AI Jev: Typed Decisions, Calibrated Confidence, and Speculative Fan-Out with a System One Model
AI & Technology

A Coding Guide to TypeSafe AI Jev: Typed Decisions, Calibrated Confidence, and Speculative Fan-Out with a System One Model

September 24, 2026
Everything Announced At Meta Connect 2026
AI & Technology

Everything Announced At Meta Connect 2026

September 24, 2026
Meta Put Muse In A Tamagotchi Like ‘Charm’ Device
AI & Technology

Meta Put Muse In A Tamagotchi Like ‘Charm’ Device

September 24, 2026
Meta Brings FDA-Cleared Hearing Enhancement To Its Smart Glasses
AI & Technology

Meta Brings FDA-Cleared Hearing Enhancement To Its Smart Glasses

September 23, 2026
Next Post
Floor & Decor Holdings: A Stock That Has Been On Our Shopping List

Floor & Decor Holdings: A Stock That Has Been On Our Shopping List

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Trump announces new U.S. Space Academy

Trump announces new U.S. Space Academy

September 22, 2026
Kalshi permanently bans George Santos

Kalshi permanently bans George Santos

September 20, 2026
Hayden Panettiere died from an overdose involving fentanyl. She had just gone to rehab, coroner says – AP News

Hayden Panettiere died from an overdose involving fentanyl. She had just gone to rehab, coroner says – AP News

September 22, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!