FTFuture Technology
HARDWARE

NVIDIA's Vera Rubin Is in Full Production and Partners Are Racing to Scale

· 3 min read · By Nath Connell

Key takeaways

  • Vera Rubin NVL72 packs 72 Rubin GPUs per rack-scale system connected via NVLink high-bandwidth interconnect
  • Production deployments are confirmed at CoreWeave and Google as publicly named partners
  • NVIDIA claims Vera Rubin delivers the lowest token cost of any system it has shipped, a key metric for AI cloud economics
  • NVIDIA describes production volumes as going gigascale, indicating a faster ramp than previous Blackwell generation supply constraints

NVIDIA's Vera Rubin architecture has moved from announcement to full production, and the numbers coming out of early partner deployments are significant enough that the rest of the AI infrastructure industry needs to take notice. The Vera Rubin NVL72 is now running in production environments at partners including CoreWeave and Google, and the performance-per-watt improvements over the previous generation are substantial.

The headline claim from NVIDIA is that Vera Rubin delivers the lowest token cost of any system they have shipped. That is a specific, measurable claim that matters enormously to the economics of running large language models at scale. Token cost is what cloud providers and enterprise AI teams actually care about when they are choosing hardware. It is not about raw performance in isolation. It is about how much compute you get per dollar of electricity and capital expenditure.

What Vera Rubin Actually Is

Vera Rubin is NVIDIA's successor to the Hopper and Blackwell architectures. The NVL72 configuration packs 72 Rubin GPUs into a single rack-scale system connected via NVLink, NVIDIA's high-bandwidth interconnect. The Vera CPU, which NVIDIA developed in-house rather than sourcing from ARM or Intel, handles system management and data processing tasks, freeing the Rubin GPUs to focus entirely on AI compute.

The NVLink interconnect in NVL72 is critical. One of the fundamental constraints on large model training and inference is the bandwidth between GPUs. If your GPUs cannot pass data to each other fast enough, they spend time waiting rather than computing, and your utilisation numbers suffer. NVIDIA's interconnect bandwidth has consistently been a competitive moat, and Vera Rubin maintains that advantage over disaggregated GPU clusters from competitors.

The performance-per-watt improvement is also timely. Data centre power consumption has become a genuine constraint on AI deployment. Hyperscalers are signing power purchase agreements that were unthinkable five years ago, and any efficiency gain at the chip level has a multiplied effect on the total cost of running an AI infrastructure. If Vera Rubin genuinely delivers meaningfully better performance per watt than Blackwell, that is not just a spec sheet win. It changes the maths on how much compute you can fit into a given power envelope.

The future, in 3 minutes a day. The biggest tech story explained every morning, free. Get the briefing →

CoreWeave and Google as Early Validators

The choice of CoreWeave and Google as the publicly named early production partners is worth noting. CoreWeave has built its entire business around being the most capable NVIDIA-based cloud provider, and it has a strong incentive to deploy the best available hardware as quickly as possible. Its validation of Vera Rubin's production readiness carries weight with the enterprise AI market that pays premium prices for reliable, high-performance compute.

Google is a more complex case. Google has its own TPU infrastructure, which it uses extensively for internal workloads and Google Cloud offerings. The fact that Google is also running Vera Rubin NVL72 in production suggests that even a company with its own world-class AI hardware has found cases where NVIDIA's ecosystem is the more efficient choice. That is a meaningful signal about where the ecosystem advantage sits.

The Gigascale Ramp

NVIDIA describes Vera Rubin as going gigascale, meaning production volumes are ramping at a rate that supports deployment across multiple major cloud providers and enterprise customers simultaneously. This is different from the early Blackwell deployments, which experienced supply constraints that created waiting lists and gave competitors a window to position alternatives.

A smoother production ramp for Vera Rubin would be strategically important. Supply constraints are one of the few structural vulnerabilities in NVIDIA's current market position. When customers cannot get the hardware they want, they evaluate alternatives. A gigascale ramp that keeps pace with demand closes that window.

The combination of performance-per-watt improvements, lowest token cost positioning, and a smoother production ramp than previous generations suggests that Vera Rubin is the most commercially important NVIDIA architecture to date. The B200 and GB200 Blackwell systems were transformative. Vera Rubin appears to be doubling down on that momentum rather than consolidating it.

Get the briefing, free

The biggest tech story, explained in 3 minutes every weekday. Choose your briefings →

Free. No spam. Unsubscribe in one click.