NVIDIA Vera Rubin Is Here and the Numbers Are Striking
Key takeaways
- Vera Rubin NVL72 packs 72 Rubin GPUs into a single rack-scale system and is now in production at partners including CoreWeave and Google
- NVIDIA is positioning the architecture specifically around performance per watt and lowest token cost, targeting inference economics
- The architecture succeeds Blackwell and is named after astronomer Vera Rubin, whose dark matter research was long overlooked
NVIDIA's Vera Rubin architecture has moved from announcement to production ramp, and the early performance claims being shared with partners paint a picture of a significant generational leap, particularly on the metrics that actually determine how much it costs to run AI at scale.
The Vera Rubin NVL72 configuration, which packages 72 Rubin GPUs into a single rack-scale system, is now in production and running at partner sites including CoreWeave and Google. NVIDIA is positioning Vera Rubin explicitly around performance per watt and token cost, the two numbers that cloud providers and enterprise AI teams care about most when they are deciding how to deploy inference workloads.
What Vera Rubin actually delivers
The architecture succeeds Blackwell, and the naming pays tribute to Vera Rubin, the astronomer whose work on galaxy rotation curves provided some of the strongest evidence for dark matter. Naming a GPU after a scientist whose contributions were overlooked for decades before being recognised says something about the cultural moment we are in, though the chip itself is less poetic and more focused on raw throughput.
The NVL72 form factor is important because it reflects how modern AI compute is actually consumed. Workloads at hyperscale do not run on individual GPUs. They run across enormous racks where networking, memory bandwidth, and inter-chip communication are as important as raw compute. NVIDIA has invested heavily in NVLink and NVSwitch interconnects within the NVL72 chassis specifically to keep data moving fast enough that the GPUs are never waiting on each other.
On token cost, which is the measure of how much it costs to process a given amount of AI output, Vera Rubin is reportedly offering meaningful reductions compared to Blackwell configurations at equivalent workloads. The exact percentages have not been published in full, but partners with early access have been vocal about the economics being favourable.
The partner ramp matters
Production ramping at CoreWeave and Google simultaneously is a sign that NVIDIA's manufacturing relationships, including the SK Hynix HBM supply discussed elsewhere today, are holding together. One of the persistent risks in high-end GPU production is that yield issues or memory supply constraints can throttle how quickly new architectures reach meaningful scale. The fact that two major cloud providers are running NVL72 racks suggests the ramp is on track.
CoreWeave's involvement is particularly telling. The company has built its entire business model around offering NVIDIA GPU capacity to AI developers at scale, and its willingness to deploy Vera Rubin early is both a vote of confidence in the architecture and a commercial necessity, since its customers will migrate to the newest hardware as quickly as it becomes available.
Google's position is more nuanced. Google builds its own AI accelerators under the TPU brand, and those chips handle a significant portion of Google's internal training workloads. But for certain inference tasks and for customers who specifically need NVIDIA compatibility, running Vera Rubin capacity through Google Cloud gives Google a competitive card to play against AWS and Azure.
What this means for AI costs
The consistent theme across the Blackwell to Vera Rubin transition is that NVIDIA is working hard to bring down the per-token cost of AI inference. This matters because inference, running a trained model to produce outputs, has become a larger share of total AI compute spend than training as more applications move into production.
If Vera Rubin genuinely delivers lower token costs at scale, it will put downward pressure on what cloud providers charge for AI API calls. That is good news for developers building products on top of AI. It also means NVIDIA is trying to expand the market by making AI compute accessible to a wider range of use cases rather than just the most expensive frontier model work.