FTFuture Technology
HARDWARE

NVIDIA Vera Rubin Is Here and It Is Already Running at Scale

· 3 min read · By Nath Connell

Key takeaways

  • Vera Rubin NVL72 production is ramping with racks live at CoreWeave and Google
  • Japan's national AI infrastructure deployment uses 13,750 Vera CPUs and 27,500 Rubin GPUs
  • NVIDIA is positioning Vera Rubin around performance per watt and lowest token cost rather than peak FLOPS
  • The NVL72 is a rack-scale system integrating 72 Rubin GPUs with NVIDIA's in-house Vera CPU

There is a moment in any hardware generation where the product stops being a press release and starts being real infrastructure. For NVIDIA's Vera Rubin architecture, that moment appears to have arrived. Production of the Vera Rubin NVL72 is ramping up, with racks already running at partners including CoreWeave and Google, and the headline number NVIDIA is leading with is performance per watt combined with the lowest token cost of any system its partners have deployed.

This matters more than it might sound. The AI industry has spent the last two years in a raw performance arms race, throwing power and money at clusters to get better benchmark numbers. But as the market matures and enterprises actually try to run AI workloads at scale without exploding their energy bills, efficiency becomes the real competitive edge. NVIDIA seems to know this, and Vera Rubin is being positioned squarely around that pitch.

What Is Actually Inside Vera Rubin NVL72

The NVL72 form factor puts 72 Rubin GPUs into a single rack-scale system, paired with NVIDIA's Vera CPUs. The Vera CPU is significant in its own right: NVIDIA designing its own CPU architecture specifically for AI factory workloads means the company is no longer just selling GPUs into someone else's server design. It is selling the whole compute stack, top to bottom.

Japan's national AI infrastructure programme gives a sense of the scale NVIDIA is targeting. That deployment alone involves 13,750 Vera CPUs and 27,500 Rubin GPUs, which is a frankly enormous single commitment to a single architecture. When a national government bets its foundational AI compute on your next-generation silicon, you have cleared a threshold most hardware companies never reach.

For CoreWeave and Google, the commercial calculus is simpler but equally telling. Both companies need to offer competitive inference pricing to win enterprise contracts. If Vera Rubin genuinely delivers lower cost per token than the H100 or Blackwell generation, cloud providers running it will be able to undercut rivals, which means pressure across the entire market to upgrade faster than planned.

Why Performance Per Watt Is the New Benchmark That Matters

Data centre power is genuinely becoming the binding constraint on AI expansion. Regulators in several countries are scrutinising energy use, utilities are struggling to provision new capacity quickly enough, and CFOs are increasingly asking hard questions about electricity costs relative to AI ROI.

The future, in 3 minutes a day. The biggest tech story explained every morning, free. Get the briefing →

NVIDIA's decision to lead the Vera Rubin launch narrative with efficiency rather than raw FLOPS is a smart read of where the industry is heading. It is also a subtle signal to enterprises that have been waiting to see whether AI infrastructure costs would ever come down: the answer, NVIDIA is suggesting, is yes, and the path runs through this generation of hardware.

The NVL72 rack-scale design also reflects a broader shift in how high-performance compute is being sold. Individual GPU boards are increasingly giving way to complete rack systems where networking, cooling, and compute are integrated from the start. This reduces deployment complexity for cloud providers but also deepens vendor lock-in, which is exactly where NVIDIA wants to be.

What This Means for the Competition

AMD's MI350 and Intel's Gaudi 3 are both chasing this market, and both have made genuine efficiency improvements. But neither has the software ecosystem depth that CUDA and NVIDIA's broader toolchain provide. Switching costs are real, and they compound over time as teams build workflows on NVIDIA's stack.

The real wildcard is custom silicon. Google's TPUs, Amazon's Trainium, and Microsoft's Maia chips are all designed to reduce dependence on NVIDIA for specific workloads. But none of them are general-purpose in the way Vera Rubin is, and none of them are available to third-party cloud providers trying to compete with the hyperscalers.

For now, the ramp of Vera Rubin NVL72 production with major cloud partners is a strong signal that NVIDIA's hardware lead is not evaporating anytime soon. The efficiency story gives it a new dimension beyond raw speed, and that is exactly what the next phase of AI infrastructure build-out is going to need.

Get the briefing, free

The biggest tech story, explained in 3 minutes every weekday. Choose your briefings →

Free. No spam. Unsubscribe in one click.