FTFuture Technology
HARDWARE

Vera Rubin Is Ramping Fast: What the Performance Per Watt Numbers Mean for Cloud AI

· 3 min read · By Nath Connell

Key takeaways

  • Vera Rubin NVL72 racks are in production at CoreWeave, Google, and other cloud partners as of July 2026
  • NVIDIA is claiming Vera Rubin delivers the lowest cost per token of any production AI infrastructure currently available
  • Performance per watt is increasingly critical as data centre power constraints limit AI infrastructure expansion globally
  • The shift from training to inference economics means token cost, not raw FLOP count, is becoming the defining commercial metric

NVIDIA's Vera Rubin architecture is past the announcement stage and into genuine production scale, and the numbers being reported by early partners are worth paying close attention to. The key metric NVIDIA is emphasising is not raw compute, but performance per watt and cost per token, which tells you a lot about where the AI infrastructure economics are heading.

Vera Rubin NVL72 racks are now running at CoreWeave, Google, and other major cloud partners, with production volume ramping according to NVIDIA's most recent update. The company is highlighting that Vera Rubin delivers what it describes as the lowest token cost currently available in production AI infrastructure, a claim that, if accurate, has significant commercial implications for every company running large-scale AI inference.

Why Token Cost Is the Number That Matters

For the past few years, AI infrastructure conversations have been dominated by training: how much compute do you need to train a frontier model, and how long does it take. But the economics of AI are shifting. Training a frontier model is something that happens once, or a handful of times per year, at immense cost. Inference, serving that model to millions or billions of users, happens continuously and compounds in cost at scale.

The cost of generating a single output token from a large language model, which is roughly equivalent to a word fragment, is vanishingly small in isolation. But multiply that by billions of daily queries across an AI-powered product, and token cost becomes one of the most important numbers in the business.

A meaningful reduction in cost per token does not just make existing AI products cheaper to operate. It makes previously uneconomical AI applications viable. Applications that require very long context windows, or that call AI models many times per user session, or that serve markets where the revenue per query is very low, all become more feasible when token cost drops significantly.

The Performance Per Watt Story

The other metric NVIDIA is pushing is performance per watt, which matters for reasons that go beyond operating costs. Data centres are running into power constraints. In many markets, the limiting factor for AI infrastructure expansion is not hardware supply but access to electrical power and the cooling systems to manage the heat that power generates.

The future, in 3 minutes a day. The biggest tech story explained every morning, free. Get the briefing →

A generation of AI hardware that delivers substantially more inference compute per watt consumed means that the same power budget can serve significantly more AI queries. For hyperscalers already negotiating with electricity grids and local governments about new data centre capacity, this is not a marginal improvement. It is the difference between expanding and waiting.

The specific efficiency gains NVIDIA is claiming for Vera Rubin over its predecessor Blackwell architecture have not been fully detailed in public benchmarks yet, but early partner reports suggest meaningful improvements on inference workloads. CoreWeave, which has a front-row seat as one of the largest Vera Rubin deployments, has been publicly positive about the efficiency profile.

What This Means for the Competitive Landscape

The infrastructure conversation in AI has been complicated by the emergence of serious alternatives to NVIDIA's stack. AMD's MI series has made genuine progress. Google's TPU v6 is impressive for Google's own workloads. Groq's inference hardware offers extraordinary speed on certain tasks. And the various custom ASIC efforts at Amazon and Microsoft are maturing.

Against this backdrop, NVIDIA's ability to demonstrate that its leading-edge architecture also leads on the metrics that matter most for commercial deployment, token cost and performance per watt, is important. It means that even customers who are exploring alternatives have a high bar to clear before switching makes financial sense.

The production ramp also matters. Having Vera Rubin actually running at scale at CoreWeave and Google is a very different position from having impressive benchmark results in a controlled environment. Real production deployments surface real problems, and the fact that NVIDIA is pointing to these partner deployments as reference cases suggests confidence that the hardware is performing as specified under real conditions.

For anyone building or buying AI infrastructure in the second half of 2026, the Vera Rubin numbers are going to be the baseline everything else gets measured against.

Get the briefing, free

The biggest tech story, explained in 3 minutes every weekday. Choose your briefings →

Free. No spam. Unsubscribe in one click.