FTFuture Technology
HARDWARE

NVIDIA's Vera Rubin Is Shipping, and the Performance Per Watt Numbers Are Striking

· 3 min read · By Nath Connell

Key takeaways

  • Vera Rubin NVL72 packs 72 Rubin GPUs into a single rack connected via NVLink for near-seamless chip-to-chip communication
  • Production racks are already running at CoreWeave and Google, with NVIDIA claiming gigascale production volumes
  • NVIDIA positions Vera Rubin as delivering the lowest token cost worldwide, targeting cloud provider procurement decisions
  • Performance per watt improvements directly address data centre power consumption constraints affecting AI deployment in the UK, Europe, and US

NVIDIA's next chip generation is no longer a roadmap slide. Vera Rubin NVL72 racks are in production and running at partners including CoreWeave and Google, and the early numbers NVIDIA is putting out around performance per watt and token cost are the kind of thing that will make cloud providers rethink their infrastructure plans for the next three years.

Vera Rubin has been NVIDIA's most anticipated architectural update since Blackwell, and arguably more important. Blackwell was a substantial generational leap, but it arrived into a market where the bottlenecks were becoming increasingly obvious: power consumption per rack was climbing fast, and the cost of generating a single AI inference token, when you factored in electricity and cooling, was stubbornly high. Vera Rubin is NVIDIA's answer to both problems simultaneously.

What Vera Rubin Actually Changes

The NVL72 configuration puts 72 Rubin GPUs into a single rack, connected via NVLink at bandwidth rates that make the chip-to-chip communication effectively seamless for large model inference. The key architectural difference from Blackwell is in how the memory subsystem and compute pipeline are integrated. NVIDIA has been tight on exact specifications in public announcements, but partners running early production racks have indicated that inference throughput per watt is meaningfully better than the equivalent Blackwell configuration.

Token cost matters enormously here. For cloud providers selling AI inference as a service, the cost of producing one million output tokens is the fundamental unit economics figure that determines whether their business model works. If Vera Rubin genuinely delivers lower token cost than Blackwell, every hyperscaler running inference at scale has a strong financial reason to accelerate their Vera Rubin procurement. That creates the demand pull that makes NVIDIA's production ramp self-reinforcing.

Who Is Running It First

CoreWeave and Google being named as early production partners is significant but not surprising. CoreWeave has built its entire business model around being the fastest to market with new NVIDIA hardware, and it has the direct supply relationships to get early allocation. Google's inclusion is more interesting: Google has its own TPU silicon and has been reducing its dependence on NVIDIA for years. The fact that Google is also running Vera Rubin suggests either that the economics are compelling enough to use both, or that Google's TPU capacity is not sufficient for every workload type it needs to serve.

The future, in 3 minutes a day. The biggest tech story explained every morning, free. Get the briefing →

The term NVIDIA is using is gigascale, which implies production volumes well beyond what they achieved in early Blackwell ramp. NVIDIA has been investing heavily in its manufacturing partnerships with TSMC, and the Vera Rubin die is manufactured on TSMC's most advanced node. Getting to gigascale quickly would require everything in that supply chain to be working smoothly, which is itself a notable operational achievement.

The Efficiency Story Is the Real Story

There is a tendency in chip coverage to focus on raw performance numbers: flops, bandwidth, clock speeds. Vera Rubin's most important story is probably the efficiency one. Data centre power consumption has become a genuine constraint on AI deployment. In the UK and Europe, planning permission for new data centres is increasingly contested because of power grid impact. In the US, utilities are scrambling to add generation capacity. Any chip that can deliver meaningfully more AI compute per watt directly addresses that constraint.

NVIDIA describing Vera Rubin as driving the lowest token cost for partners worldwide is a specific, commercial claim aimed directly at CFOs and infrastructure procurement teams, not just engineers. It is NVIDIA positioning Vera Rubin not just as faster but as the economically correct choice, which is a different and arguably more powerful argument.

The ramp is underway. Production is live. The real test will be whether those early efficiency numbers hold up at the scale of millions of chips deployed across dozens of data centres running thousands of different model types. But based on what is shipping today, Vera Rubin looks like the chip that will define AI infrastructure economics for the next two to three years.

Get the briefing, free

The biggest tech story, explained in 3 minutes every weekday. Choose your briefings →

Free. No spam. Unsubscribe in one click.