NVIDIA's Vera Rubin Is Already Delivering the Lowest Token Costs in the Industry
Key takeaways
- Vera Rubin NVL72 racks are running at production scale with CoreWeave and Google
- The NVL72 form factor connects 72 Rubin GPUs via NVLink in a single rack-scale system
- NVIDIA's primary claim is best-in-class performance per watt and lowest token cost for cloud inference
- Power efficiency is increasingly critical as data centres hit grid capacity limits in most major markets
When NVIDIA announced the Vera Rubin architecture, the promise was straightforward: more performance per watt at a lower cost per token than anything before it. Now that production is actually ramping, we're starting to see whether that promise holds up in the real world. Spoiler: it largely does, and the implications for how AI inference gets priced are significant.
Vera Rubin NVL72 racks are now running at scale with partners including CoreWeave and Google. That's not a small-scale pilot. These are two of the largest AI infrastructure operators on the planet, and the fact that they're both live with Vera Rubin gives the performance claims considerably more credibility than a press release alone would.
What the Numbers Actually Mean
The headline metric NVIDIA is pushing is performance per watt. This matters more than raw throughput for most operators, because electricity is now one of the dominant costs in running AI at scale. Data centres are bumping up against power limits in almost every major market. A chip that delivers more useful compute per watt consumed doesn't just save money on energy bills; it also lets operators run more capacity within existing infrastructure constraints, which is a very big deal when new grid connections can take years to arrange.
Lower token cost is the downstream effect. When cloud providers can serve more tokens per kilowatt-hour, they can price AI APIs more competitively, which in turn drives adoption among developers and enterprises. The whole ecosystem benefits when inference gets cheaper, and Vera Rubin appears to be pushing that in the right direction.
The NVL72 form factor is worth understanding. It packs 72 Rubin GPUs into a single rack-scale system, with NVIDIA's NVLink fabric connecting them at very high bandwidth. This means large models can be split across the rack without the kind of communication bottlenecks that slow down inter-GPU workloads. For frontier models that no longer fit on a single GPU or even a single node, this architecture is particularly well suited.
The Competitive Context
Vera Rubin arrives at a moment when competition in the AI chip market is genuinely heating up. AMD's MI400 series, Google's TPU v6, and a growing cohort of custom silicon from hyperscalers are all vying for the workloads that currently run on NVIDIA hardware. NVIDIA's moat has historically been the CUDA software ecosystem as much as the hardware itself, and Vera Rubin doesn't change that dynamic. But the hardware improvements do give partners a clear performance upgrade path that keeps them invested in the NVIDIA stack.
For CoreWeave specifically, Vera Rubin is strategically important. The company has built its entire business on renting NVIDIA GPU capacity to AI developers, and being an early Vera Rubin partner puts it ahead of competitors who won't have access to the same hardware for months. Google's participation is slightly different in character; Google has its own TPUs and doesn't depend on NVIDIA the way CoreWeave does, but running Vera Rubin in parallel suggests the hardware is genuinely competitive even for a company with world-class custom silicon.
What This Means for AI Pricing
The practical effect for anyone buying AI API access is that prices should continue to fall. Token costs have dropped dramatically over the past two years as hardware improved and competition intensified, and Vera Rubin is another nudge in that direction. Whether those savings flow through to end users quickly depends on market dynamics, but the underlying economics are moving in the right direction.
For enterprises evaluating whether to run AI on-premises or in the cloud, Vera Rubin's efficiency improvements make the cloud argument stronger, at least for inference workloads. On-premises deployments are still compelling for data privacy reasons, but the cost-per-query gap between cloud and self-hosted is likely to widen as next-generation silicon rolls out through hyperscaler fleets faster than enterprise buyers can refresh their own hardware.
The ramp is still early. Full production volumes take time, and the transition from announcement to widespread availability has historically taken longer than NVIDIA's optimistic timelines suggest. But the signal from CoreWeave and Google running live workloads is meaningful. Vera Rubin is real, it's shipping, and the performance numbers appear to be holding up under production conditions.