• bitcoinBitcoin(BTC)$84,495.00-0.26%
  • ethereumEthereum(ETH)$2,664.55-1.32%
  • tetherTether(USDT)$1.000.02%
  • binancecoinBNB(BNB)$766.43-0.66%
  • rippleXRP(XRP)$1.48-1.25%
  • usd-coinUSDC(USDC)$1.000.01%
  • solanaSolana(SOL)$118.160.06%
  • tronTRON(TRX)$0.334291-0.32%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.052.76%
  • zcashZcash(ZEC)$1,288.08-2.96%
  • HyperliquidHyperliquid(HYPE)$86.30-0.78%
  • dogecoinDogecoin(DOGE)$0.092127-2.35%
  • chainlinkChainlink(LINK)$13.68-4.47%
  • moneroMonero(XMR)$539.97-0.78%
  • whitebitWhiteBIT Coin(WBT)$84.06-0.62%
  • USDSUSDS(USDS)$1.000.02%
  • cardanoCardano(ADA)$0.241380-2.11%
  • leo-tokenLEO Token(LEO)$8.991.18%
  • RainRain(RAIN)$0.011144-7.86%
  • stellarStellar(XLM)$0.213582-2.64%
  • bitcoin-cashBitcoin Cash(BCH)$307.95-0.45%
  • nearNEAR Protocol(NEAR)$4.66-3.15%
  • uniswapUniswap(UNI)$8.80-3.29%
  • litecoinLitecoin(LTC)$69.841.89%
  • Ethena USDeEthena USDe(USDE)$1.000.01%
  • avalanche-2Avalanche(AVAX)$10.67-2.88%
  • CantonCanton(CC)$0.118693-1.07%
  • Blockchain USDBlockchain USD(USDB)$0.871,000.00%
  • suiSui(SUI)$1.13-4.18%
  • daiDai(DAI)$1.00-0.01%
  • hedera-hashgraphHedera(HBAR)$0.100632-1.73%
  • USD1USD1(USD1)$1.000.01%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.49-5.20%
  • BitwayBitway(BTW)$1.38-3.45%
  • tether-goldTether Gold(XAUT)$4,142.80-0.80%
  • quant-networkQuant(QNT)$230.87-5.76%
  • shiba-inuShiba Inu(SHIB)$0.000006-2.04%
  • crypto-com-chainCronos(CRO)$0.066143-3.42%
  • BittensorBittensor(TAO)$287.78-5.04%
  • Global DollarGlobal Dollar(USDG)$1.00-0.01%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.01%
  • aaveAave(AAVE)$181.265.67%
  • okbOKB(OKB)$120.34-0.51%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • Pump.funPump.fun(PUMP)$0.005282-8.28%
  • MemeCoreMemeCore(M)$1.060.76%
  • Circle USYCCircle USYC(USYC)$1.140.03%
  • OndoOndo(ONDO)$0.491168-0.20%
  • EthenaEthena(ENA)$0.231664-5.16%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.150.10%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

NVIDIA Announces DGX Spark 64GB: A 1-PetaFLOP Grace Blackwell Desktop for Local AI Agents, Fine-Tuning, and Inference

October 2, 2026
in AI & Technology
Reading Time: 10 mins read
A A
NVIDIA Announces DGX Spark 64GB: A 1-PetaFLOP Grace Blackwell Desktop for Local AI Agents, Fine-Tuning, and Inference
ShareShareShareShareShare

NVIDIA announced a new 64GB configuration of DGX Spark — from Acer, ASUS, Dell, Gigabyte, HP and MSI — its GB10-powered desktop AI system. It gives developers a way to start with one system for local models and agents, then cluster two 64GB units for 128GB of memory across the cluster and more compute when workloads grow.

The direct message is: Run open models and always-on agents on your own desk, then cluster DGX Spark system as the work grows – instead of on a metered API.

YOU MAY ALSO LIKE

Bloomberg Tech Screentime Special

AI Moves From Hollywood Tool to Hollywood Studio

The timing is not accidental. Agents burn tokens continuously through tool calls, retries, long context, and multi-step plans. Token consumption has grown 14x since early 2026. On a cloud API, every one of those tokens is billed. On owned hardware, there is no per-token fee.

The 64GB model keeps the same GB10 Grace Blackwell superchip, NVIDIA CUDA accelerated AI software stack, and ConnectX-7 networking as the original — updated with 64GB of unified LPDDR5x instead of 128GB. NVIDIA positions that as enough for today’s most capable 30–35B class open models. NVIDIA continues to offer the 128GB DGX Spark for larger single-box workloads, and clustering 64GB units adds memory and compute together.

What is Inside the Box

GB10 pairs a Blackwell GPU with 5th-generation Tensor Cores and a 20-core Grace Arm CPU. The GPU delivers up to 1 petaFLOP of FP4 AI compute, with sparsity.

Spec DGX Spark 64GB
Superchip NVIDIA GB10 Grace Blackwell
CPU 20-core Arm (10× Cortex-X925 + 10× Cortex-A725)
AI compute Up to 1 petaFLOP FP4 (with sparsity)
Memory 64GB LPDDR5x, coherent unified
Memory bandwidth 273 GB/s
Storage 1, 2, or 4TB NVMe M.2, self-encrypting
Networking ConnectX-7 NIC, 200GbE; Wi-Fi 7, BT 5.3
Display 1× HDMI 2.1a
OS NVIDIA DGX OS (Ubuntu-based)
Size / weight 150 × 150 × 50.5 mm / 1.2 kg
Max local model size Up to 100B parameters
Availability Oct 23, 2026 — NVIDIA Marketplace, OEM partners, retail

Source: NVIDIA specifications.

Unified memory is the key design choice: The CPU and GPU share one pool over NVLink-C2C, at 5x the bandwidth of PCIe Gen 5. There is no copying weights between system RAM and VRAM. For agents, that means several models, their KV caches, and tool processes live in one address space.

The software is ready on first boot: DGX OS ships with the NVIDIA AI stack, including PyTorch, Jupyter, and Ollama. NVIDIA NemoClaw installs with a single command. It adds privacy and security controls to OpenClaw agents. NVIDIA OpenShellTM, part of the NVIDIA Agent ToolkitTM, adds policy-based guardrails on top. NVIDIA NemotronTM models are optimized for the box.

It runs on a standard wall outlet: No server room, no special cooling. That is important for an agent meant to run around the clock.

Built-in networking for clustering: ConnectX-7 lets two DGX Spark 64GB systems cluster for 128GB of memory and more compute. NVIDIA Sync Cluster Assistant simplifies setup.

Which Open Models Fit in 64GB

Hardware is half the story. The other half is that 30B-class open models got good enough for agent work.

Model Developer Type Footprint Role on a Spark
Muse Glimmer Meta 29.6B dense, text + image, Apache 2.0 ~17GB (quantized) Main agent model
Nemotron 3.5 Lightning NVIDIA 30B MoE (30B-A3B) NVFP4 checkpoint Fast executor for long-running agents
Qwen3.8-27B Alibaba Qwen 27B dense ~13.5GB weights (4-bit) General agent and coding

Footprints are weight-only estimates; KV cache and runtime overhead come on top.

Muse Glimmer is the main model. Meta distilled it from Muse Spark, the model family behind the Meta AI assistant. It targets local agents: reliable tool calls, long multi-step tasks, and recovery from failures. Meta reports 51.2 on SWE-Bench Pro and 75.5 on MCP Atlas. Context runs to 131K tokens. It is also available as an NVIDIA NIM.

Full BF16 Glimmer needs 55GB+, which leaves almost nothing for context. The ~17GB quantized build is the practical choice on 64GB.

Larger models like DeepSeek V4 Flash need more than one box. NVIDIA’s own benchmarks run it on 4 64 GB clustered Sparks.

5 Things You Can Build on One Box

1. An always-on personal agent

Install Hermes and point it at Muse Glimmer or Nemotron 3.5 Lightning. Give it tools: your GitHub repos, a test runner, an RSS feed of arXiv categories. Let it run overnight.

By morning it has triaged new issues, reproduced a failing test, and drafted a pull request for review. It has also summarized the 30 papers you would never have opened. Your private notes, code, and email never leave the machine.

2. Fine-tune a coding model on your own repo

QLoRA on a 70B model fits in 64GB. Train it on your codebase, internal docs, and past PR reviews. Then serve it locally as a coding assistant that knows your conventions.

NVIDIA measured ~18,400 tokens/s on a single node for nanochat distributed fine-tuning. 

3. A day-1 model evaluation bench

A new open-weight model drops on Hugging Face. Pull it through Ollama or vLLM the same day. Run your own question set against it.

Score accuracy, then measure time to first token (TTFT) and tokens per second. Compare it against your current model. There is no API bill for re-running the eval 50 times.

4. A multi-model agent team

Unified memory lets several models share one pool. Run Bonsai 2 as a router, and Glimmer as the main reasoning agent. That is roughly 23GB of weights.

The rest covers KV cache, the OS, and tool processes. When a problem exceeds local capacity, the agent can send a sanitized question to a larger cloud model.

5. Edge and robotics prototyping

Fine-tune a vision transformer for a specific task, like spotting anomalies on a factory floor camera. Validate it locally with the same CUDA stack. Then deploy it to an NVIDIA Jetson device at the edge.

When One Box is Not Enough: Clustering with NVIDIA Sync

Every DGX Spark ships with ConnectX-7 at 200GbE. Clustering is a native feature, not an add-on.

Setup Pooled memory AI compute (FP4) Connection
1 Spark (64GB) 64GB Up to 1 PFLOP —
2 Sparks (64GB each) 128GB Up to 2 PFLOPS Direct QSFP cable, no switch
2 Sparks (128GB each) 256GB Up to 2 PFLOPS Direct QSFP cable, no switch
3 Sparks (128GB each) 384GB Up to 3 PFLOPS QSFP ring, no switch
4 Sparks (128GB each) 512GB Up to 4 PFLOPS 200GbE switch (QSFP56-DD, RoCE v2)

Source: NVIDIA. 3-node compute derived from per-unit specs.

NVIDIA states 2 clustered 64GB units deliver up to 1.7x the performance of one 128GB DGX Spark. The reason is doubled AI compute and bandwidth: 2 units provide up to 546 GB/s combined, double a single box.

NVIDIA Sync is the management layer. The Windows and macOS app discovers Sparks on your network and manages SSH access. Its Cluster Assistant configures ConnectX-7 networking for up to 4 systems. Nodes can also connect across locations over a Tailscale mesh, with no cloud in the data path.

The pattern is quite clear. Clustering roughly halves time to first token per doubling and scales fine-tuning near-linearly. Decode improves more modestly, about 1.4x at 4 nodes. For agents that read long inputs, the TTFT gain is the one that matters.

Step-by-step multi-node guides, including vLLM on stacked Sparks, are at build.nvidia.com/spark.

What It is, and What It is Not

  • Not a chat server for 100 users: 273 GB/s of bandwidth limits concurrent decode throughput.
  • Built for long-input, short-output work: Reading a repo, a log dump, or a paper stack, then writing a short result.
  • 64GB caps one box at 100B parameters: For more, cluster 2 units or choose the 128GB configuration.

Key Takeaways

  • DGX Spark 64GB keeps GB10, 1 PFLOP FP4, and the full NVIDIA AI stack.
  • 64GB runs today’s 30–35B class open models, like Qwen 3.8 27B and Nemotron 3.5 Lightning.
  • Best uses: always-on agents, QLoRA fine-tuning, and day-1 model evals — no per-token fees.
  • 2 units cluster over ConnectX-7 via NVIDIA Sync for 128GB and up to 1.7x a 128GB Spark.
  • Available October 23; the DGX Spark 128GB remains available.

Check out the DGX Spark product page, NVIDIA Sync, and DGX Spark playbooks.


Thanks to the NVIDIA team for the thought leadership / resources for this article. This article is sponsored by NVIDIA.


Jean-marc is a successful AI business executive .He leads and accelerates growth for AI powered solutions and started a computer vision company in 2006. He is a recognized speaker at AI conferences and has an MBA from Stanford.

Credit: Source link

ShareTweetSendSharePin

Related Posts

Bloomberg Tech Screentime Special
AI & Technology

Bloomberg Tech Screentime Special

October 2, 2026
AI Moves From Hollywood Tool to Hollywood Studio
AI & Technology

AI Moves From Hollywood Tool to Hollywood Studio

October 2, 2026
SpaceX Launches Crew for NASA on Retiring Dragon Capsule
AI & Technology

SpaceX Launches Crew for NASA on Retiring Dragon Capsule

October 2, 2026
Luma AI Bets on ‘Hybrid’ Future for Hollywood
AI & Technology

Luma AI Bets on ‘Hybrid’ Future for Hollywood

October 2, 2026
Next Post
Reporters detail witnessing failed execution of Christa Pike

Reporters detail witnessing failed execution of Christa Pike

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
US Senate blocks resolution demanding accountability over Israeli violence in West Bank – theguardian.com

US Senate blocks resolution demanding accountability over Israeli violence in West Bank – theguardian.com

September 30, 2026
Hochul appoints Letitia James as special prosecutor in Cornell sexual assault case

Hochul appoints Letitia James as special prosecutor in Cornell sexual assault case

October 2, 2026
Times Square may score another Asian entertainment lease

Times Square may score another Asian entertainment lease

September 27, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!