FTFuture Technology
HARDWARE

NVIDIA Vera Rubin Is Now Genuinely in Production and the Performance Numbers Are Real

· 3 min read · By Nath Connell

Key takeaways

  • Vera Rubin NVL72 is a 72-GPU rack system where all GPUs operate as a single unified compute unit via NVLink
  • Production deployments are now running at CoreWeave and Google
  • NVIDIA claims Vera Rubin delivers the lowest token cost of any current AI inference hardware
  • The next NVIDIA architecture after Vera Rubin is expected to be called Feynman
  • Power availability rather than GPU supply is increasingly the binding constraint on hyperscale AI expansion

NVIDIA's Vera Rubin architecture has moved from announcement to actual production deployment, and the performance-per-watt numbers being quoted are significant enough to warrant close attention from anyone who cares about AI infrastructure economics. The NVL72 rack system is now running at partners including CoreWeave and Google, and NVIDIA is positioning Vera Rubin as the platform that makes high-volume AI inference economically viable at a scale that previous generations could not achieve.

The framing around "lowest token cost" is the number that matters most in this context. Token cost is the metric AI infrastructure buyers care about most: how much does it cost to generate one thousand tokens of output? Everything else, power draw, interconnect speed, memory bandwidth, ultimately feeds into that single number. If Vera Rubin genuinely drives token costs down materially compared to H100 and H200 systems, it reshapes the economics of running large language model workloads for every company in the space.

What the NVL72 Architecture Delivers

The NVL72 is a 72-GPU rack-scale system, meaning all 72 GPUs operate as a single unified compute unit connected by NVLink. This is distinct from traditional GPU cluster approaches where communication between GPUs involves more overhead. The NVLink fabric inside the NVL72 allows GPUs to share memory and pass data between themselves at speeds that external interconnects cannot match.

For inference workloads specifically, this matters because large language models do not parallelise in the same way that training workloads do. The communication overhead between GPUs during inference can be a significant bottleneck. The NVL72's unified memory architecture reduces that overhead and allows the system to handle very large models, or very high request volumes, more efficiently than disaggregated GPU setups.

The performance-per-watt angle is also commercially significant for a less obvious reason. Hyperscalers like Google and cloud providers like CoreWeave are increasingly constrained not by how many GPUs they can buy but by how much power they can get to their data centres. In many markets, power is the binding constraint on AI infrastructure expansion. A chip that delivers meaningfully more compute per watt effectively unlocks capacity that the power grid was otherwise limiting.

The Competitive Landscape

AMD's MI350 series and Intel's Gaudi 3 accelerators are the main alternatives, and both companies have been competitive on price. But benchmark performance in controlled settings and real-world production performance at scale are different things, and NVIDIA's advantage in software, particularly the CUDA ecosystem and the optimisation work done by partners, tends to compound over time.

The future, in 3 minutes a day. The biggest tech story explained every morning, free. Get the briefing →

The fact that CoreWeave and Google are already running Vera Rubin in production is itself a signal. These are organisations that run rigorous evaluations before deploying new hardware at scale. Their adoption suggests the real-world performance matches what the spec sheet promises, at least closely enough to justify switching from previous generation hardware.

The broader trend here is acceleration rather than plateau. Each generation of NVIDIA GPU has delivered meaningful step-up in AI compute performance, and Vera Rubin continues that trajectory. The concern occasionally raised by analysts, that AI model architecture improvements might outpace the need for raw compute, has so far not materialised. Demand for compute continues to exceed supply, and the market continues to absorb each new generation of hardware at substantial scale.

What Comes After Vera Rubin

NVIDIA has already publicly indicated that its next architecture after Vera Rubin will be called Feynman, maintaining the tradition of naming GPU generations after physicists. The cadence of new architectures has been approximately annual for training-class hardware, which keeps the upgrade pressure on hyperscale buyers constant.

For enterprises and smaller cloud providers, this creates a complicated planning challenge. Hardware that is state of the art today may be a generation behind within twelve to eighteen months. The economics of when to buy, when to wait, and how to structure depreciation around rapidly evolving hardware is a genuinely difficult operational problem that the AI infrastructure boom has created for almost every organisation trying to build at scale.

Vera Rubin is the answer to the current question. The question itself, however, keeps moving.

Get the briefing, free

The biggest tech story, explained in 3 minutes every weekday. Choose your briefings →

Free. No spam. Unsubscribe in one click.