Future TechnologyFuture Technology
AI

Nvidia Killed Rubin CPX and Paid 20 Billion Dollars for SRAM Instead

· 3 min read · By Future Technology

Key takeaways

  • Rubin CPX was announced for massive-context inference but never reached production silicon
  • Nvidia replaced it with a Groq 3 LPX rack following a 20 billion dollar licensing deal
  • Groq designs keep weights in on-chip SRAM instead of streaming from HBM, trading capacity for latency

Rubin CPX was supposed to be a new class of Nvidia GPU, built specifically for massive-context inference. It never reached production silicon. Nvidia pulled it from the public roadmap at GTC 2026 and put something else in its place.

What replaced it

The slot Rubin CPX was meant to fill now belongs to the Groq 3 LPX rack, which arrived off the back of a 20 billion dollar licensing deal with Groq. The Vera Rubin rack-scale system carrying Groq 3 LPX went into full production in late August.

That is a lot of money to spend on not building something yourself.

The part that actually matters is the memory

Rubin CPX was a conventional GPU answer to long context: more compute, more high bandwidth memory, stream the weights in as you need them. Groq took a different route. Its designs are SRAM-based, holding model weights in fast on-chip memory rather than pulling them from HBM on every pass.

That trade is capacity for latency. You fit less, and what you do fit responds much faster. For training, where you are grinding through enormous batches over weeks, capacity wins comfortably. For serving tokens to people who are sitting there waiting, latency is the whole product.

Nvidia has spent a decade defining what training hardware looks like. Licensing someone else architecture for the serving side is a clear read on where the economics landed.

Why inference economics flipped

Training a frontier model is a large one-off cost. Serving it is a bill that arrives every day, forever, and scales with how many people use the thing. As deployment has broadened, inference has quietly become most of the compute spend rather than a rounding error after training.

Once that happens, the hardware question changes shape. It stops being about how fast you can move through a dataset and becomes about cost per token at an acceptable response time. Memory architecture decides that number, which is why the industry has spent 2026 arguing about SRAM, HBM and everything between them.

The capital markets have been reading the same tea leaves. The Broadcom custom silicon deal with Anthropic was structured around serving capacity, not training runs, and the Anthropic and OpenAI race to public markets is being underwritten on inference revenue.

What it means for everyone else

Nothing immediately. Nobody is buying a Vera Rubin rack for a home office, and consumer parts like the RTX Spark line run on entirely different assumptions.

The second-order effect is worth watching, though. If SRAM-heavy designs really do win the serving half at scale, that architectural argument tends to filter downward within a couple of generations. The chips in your machine in 2029 will have been shaped by a decision Nvidia made about a rack in 2026.

Nvidia has not published a technical post-mortem on why CPX stalled, and it is unlikely to. The 20 billion dollars is the statement.

Get the briefing, free

The biggest tech story, explained in 3 minutes every weekday. Choose your briefings →

Free. No spam. Unsubscribe in one click.