Vera Rubin NVL72 Is Ramping Fast and the Partners Running It Are Already Claiming Record Efficiency
Key takeaways
- Vera Rubin NVL72 packs 72 GPUs into a single rack-scale unit connected via NVLink high-bandwidth interconnect
- Partners including CoreWeave and Google are already running Vera Rubin at production scale
- NVIDIA is positioning Vera Rubin as delivering the lowest token cost of any AI compute platform currently available
- Data centre power availability, not capital or silicon, is the primary constraint on AI infrastructure growth in 2026
NVIDIA's Vera Rubin architecture has moved from announcement to serious production scale faster than almost anyone predicted, and the partners deploying it are starting to put numbers on what that actually means in practice.
Vera Rubin NVL72 production is ramping across partners including CoreWeave and Google, with NVIDIA describing the deployment as going "gigascale". That is a deliberately vague term, but it signals that Vera Rubin is no longer in early-access territory. This is mainstream infrastructure for the largest AI workloads on the planet.
What Makes Vera Rubin Different
Vera Rubin is NVIDIA's successor to the Blackwell architecture. The NVL72 configuration packs 72 GPUs into a single rack-scale unit connected by NVLink, NVIDIA's high-bandwidth interconnect. The key selling point is performance per watt: Vera Rubin is designed to deliver substantially more compute per unit of energy than its predecessors, which matters enormously when you are running infrastructure at this scale.
Data centre power is the binding constraint for AI right now. The industry is not short of capital, silicon ambition, or software talent. It is short of power. Every megawatt saved per rack translates directly into more AI capacity that can be deployed within existing grid connections and power purchase agreements. That is why NVIDIA's framing of Vera Rubin around efficiency rather than raw performance is strategically smart.
The token cost angle is equally significant. For cloud providers selling AI inference as a service, the cost per million tokens generated is the core unit economics metric. If Vera Rubin genuinely delivers the lowest token cost of any platform currently available, that is a compelling commercial proposition for every hyperscaler and AI cloud provider choosing what to build on.
The CoreWeave and Google Deployments
CoreWeave and Google are doing serious work here. CoreWeave has been one of NVIDIA's most aggressive infrastructure partners, building GPU cloud capacity at a pace that has surprised even optimistic observers. Its Vera Rubin deployment puts it in a position to offer frontier AI compute to enterprise customers who cannot or do not want to build their own infrastructure.
Google's deployment is more nuanced. Google has its own custom AI accelerator programme in TPUs, and the fact that it is also deploying Vera Rubin hardware suggests that even with substantial in-house silicon capability, the economics or performance characteristics of NVIDIA's latest architecture are compelling enough to warrant running both.
The Power Efficiency Story
NVIDIA's emphasis on performance per watt is not just marketing language. The NVL72's architecture benefits from NVLink's ability to move data between GPUs without going through the slower PCIe bus, reducing the energy overhead of inter-chip communication. Combined with advances in the underlying GPU die efficiency, this compounds into meaningful power savings at the rack level.
For hyperscalers operating tens of thousands of these systems, even a 10% improvement in performance per watt translates into hundreds of megawatts of freed capacity. Given that new data centre power connections in most major markets involve multi-year planning timelines, that freed capacity has real, immediate value.
What This Means for the Competitive Landscape
AMD's MI400 series and Intel's Gaudi successors are both targeting this market, but neither has matched NVIDIA's ecosystem depth. The software stack, the CUDA compatibility, the NVLink interconnect, and now the Omniverse and Agent Toolkit integrations all create switching costs that make Vera Rubin deployments sticky even for organisations that would prefer more vendor diversity.
The Arm-based Vera CPU, which pairs with the Rubin GPUs in the NVL72 configuration, adds another layer. NVIDIA is not just selling GPUs anymore; it is selling a complete compute architecture. That changes the competitive calculus for procurement teams who have historically been able to mix and match CPU and GPU vendors more freely.
For anyone watching the AI infrastructure market, Vera Rubin going gigascale is the hardware story of 2026. The efficiency numbers are the ones to watch as more deployment data comes in over the next two quarters.