NVIDIA Vera Rubin NVL72 Ramps Up at CoreWeave and Google
Key takeaways
- Vera Rubin NVL72 production is actively ramping with racks running at CoreWeave and Google
- NVL72 packs 72 GPUs into a rack-scale system for superior interconnect bandwidth and memory coherence
- NVIDIA positions Vera Rubin on performance per watt and lowest token cost versus the Hopper generation
- NVIDIA describes deployment ambitions as gigascale, implying millions of GPUs across global data centres
NVIDIA's next-generation Vera Rubin architecture is no longer a roadmap slide. Production of the Vera Rubin NVL72 is actively ramping, with racks already running at cloud partners CoreWeave and Google. For anyone tracking where AI compute is heading, this is one of the most significant hardware developments of 2026.
The NVL72 designation refers to a rack-scale system configuration, part of NVIDIA's shift toward thinking about compute not at the individual GPU level but at the rack level. Packing 72 GPUs into a single networked unit allows for a degree of interconnect bandwidth and memory coherence that you simply cannot achieve by stringing together individual cards. It is a fundamentally different approach to system design, and it has major implications for the kinds of AI workloads that can run efficiently on the hardware.
Performance Per Watt as the New Battleground
NVIDIA is billing Vera Rubin as a step forward in performance per watt, and that framing is telling. Raw performance numbers matter, but data centres are increasingly constrained by power and cooling rather than physical space or chip supply. A system that delivers more compute per unit of electricity is not just cheaper to run, it is often the difference between a deployment being feasible and it not being feasible at all, given that power grid capacity is becoming a genuine limiting factor in AI infrastructure expansion.
The lowest token cost claim is equally pointed. In the current market, the cost per token generated is one of the primary metrics that cloud providers and enterprise buyers use to evaluate AI infrastructure. Every fraction of a cent matters at scale when you are processing billions of queries. If Vera Rubin genuinely delivers a lower cost per token than the Hopper generation it is replacing, adoption by hyperscalers will be rapid. CoreWeave and Google are not running these racks out of curiosity.
What Gigascale Actually Means
NVIDIA used the word gigascale in its announcement, and that is worth unpacking. It suggests that the deployment ambition here is not hundreds or thousands of GPUs but millions, spread across multiple data centres and cloud regions globally. Gigascale deployments require a level of supply chain reliability, software stability, and partner ecosystem depth that only a handful of companies in the world can manage. The fact that CoreWeave, a specialist AI cloud provider, is one of the first partners listed alongside Google says something about how the AI infrastructure market has matured. CoreWeave has grown from a crypto mining operation into one of the most significant AI compute providers in the world in the space of a few years.
For NVIDIA, Vera Rubin represents the first major architecture transition since Hopper, which powered the H100 and H200 GPUs that became the defining compute substrate of the first wave of large language model deployment. Maintaining momentum through an architecture transition is always a risk, not technically but commercially. Enterprise buyers and cloud providers have existing software stacks, CUDA codebases, and operational muscle memory built around Hopper. Convincing them to migrate requires the new architecture to be substantially better, not just marginally so.
The early performance per watt and token cost messaging suggests NVIDIA is confident the numbers are compelling enough to make that migration argument straightforwardly. The ramp at CoreWeave and Google is the proof of concept. The rest of the hyperscaler community will be watching those deployments closely before committing their own purchasing decisions for 2027 and beyond.
Vera Rubin is the architecture NVIDIA hopes will anchor the next phase of AI infrastructure. If the production ramp holds and the performance claims bear out in real deployments, it is likely to do exactly that.