Future TechnologyFuture Technology
Computing

Nvidia RTX Spark: A Petaflop And 128GB Of Memory On A Desk

· 4 min read · By Future Technology

Key takeaways

  • RTX Spark pairs a 20 core Grace CPU with a Blackwell RTX GPU over NVLink-C2C for up to 1 petaflop of AI compute
  • 128GB of unified memory is the headline number, because local model work has been gated by VRAM for two years
  • Eight PC brands are signed up as OEM launch partners as of August 2026, with no confirmed USD pricing and a release window of this fall
  • Nvidia also shipped DLSS 4.5 Ray Reconstruction in August, folding denoising and super resolution into a single transformer model

Nvidia has bolted a 20 core Grace CPU to a Blackwell RTX GPU over NVLink-C2C and called the result RTX Spark. The headline specs are up to 1 petaflop of AI compute and 128GB of unified memory, with eight PC brands signed up as OEM launch partners as of August 2026.

No USD pricing has been confirmed. The release window is still just this fall.

The Nvidia RTX Spark specs that actually matter

Ignore the petaflop for a second. The number to look at is 128GB of unified memory.

Local model work has been gated by VRAM for two years. A 24GB consumer card tops out around a 30 billion parameter model at usable quantisation, which is why so many people building at home end up renting cloud time or wiring together a multi-GPU rig with all the power draw and cooling noise that implies. We went through those thresholds in detail in our guide to local LLM VRAM requirements in 2026.

128GB of unified memory moves that ceiling somewhere else entirely. Unified means the CPU and GPU address the same pool, so you are not shuttling weights across a bus and paying for it in latency. If you want the background on why memory type and bandwidth decide so much of this, HBM vs GDDR vs DDR explained covers the tradeoffs, and high bandwidth flash is the next attempt at the same problem from a different angle.

What changes if the box holds the model

The interesting shift is not raw speed, it is where the work happens.

A desk-side machine that can hold a large model in memory changes the calculus on three things at once: privacy, because the data never leaves the room; latency, because there is no round trip; and running costs, because you pay once instead of per token. For a small team handling client data or medical records, the first of those is worth more than the benchmark.

Nvidia also shipped DLSS 4.5 Ray Reconstruction in August, a second generation transformer model that folds denoising and super resolution into a single network across all GeForce RTX cards. Different problem, same pattern of pushing more work into one model.

What you can do today while you wait

No price, no date, and a spec sheet is not a product. If you want more local headroom before Spark ships, the practical routes are the ones that already exist.

A current-generation flagship card gets you 32GB of GDDR7 and a lot of compute now. The ASUS ROG Astral RTX 5090 32GB is available on Amazon, though it wants a serious power budget and a case with real airflow.

If you are running larger models slowly rather than smaller models quickly, system memory is the cheaper lever. A Corsair Vengeance DDR5 6000MHz kit is on Amazon and pairs sensibly with offloading.

And if unified memory is the specific thing you are after, Apple has been shipping it for years. The Mac mini M4 with 24GB of unified memory is on Amazon, which is a fraction of what Spark promises but available at a known price this week.

The question nobody has answered

Pricing. A petaflop and 128GB of unified memory on a desk is only interesting if it costs less than the cloud time it replaces. Until Nvidia puts a number on it, everything else is a spec sheet.

Some links in this article are affiliate links. We may earn a small commission at no extra cost to you.

Read next

Get the briefing, free

The biggest tech story, explained in 3 minutes every weekday. Choose your briefings →

Free. No spam. Unsubscribe in one click.