FTFuture Technology
AI

NVIDIA Vera Rubin Is Now in Full Production, and the Token Cost Numbers Are Striking

· 3 min read · By Nath Connell

Key takeaways

  • NVIDIA Vera Rubin NVL72 racks are in active production at CoreWeave and Google
  • The NVL72 system bundles 72 Rubin GPUs with Vera CPUs into a rack-scale unit, sold as a complete system rather than individual components
  • NVIDIA is positioning Vera Rubin as delivering the lowest cost per token of any available AI inference hardware
  • Performance per watt is the key metric as data centres face electricity supply constraints limiting AI infrastructure growth

NVIDIA's Vera Rubin architecture has moved from announcement to full-scale production, and the company is making a pointed argument: this is now the most cost-efficient way to run large language model inference at scale. With Vera Rubin NVL72 racks already running at CoreWeave and Google, the platform is no longer a roadmap promise. It's infrastructure.

The headline claim from NVIDIA is about performance per watt and token cost. In AI infrastructure, cost per token is the number that matters most to the companies actually running models commercially. Every query you send to a chatbot, every document a model processes, every code suggestion an AI IDE spits out, all of that is a token. At scale, even small reductions in cost per token translate to enormous savings.

What Vera Rubin NVL72 Actually Is

The NVL72 is a rack-scale system, not a single chip. It bundles 72 Rubin GPUs together with NVIDIA's Vera CPUs and high-bandwidth networking into a single cohesive unit that's designed to be deployed in large AI factories. The idea is that you don't buy one GPU and add it to a server; you buy a rack that's been engineered as a complete system, with all the thermal management, interconnects, and software stack pre-integrated.

This is a meaningful shift from how GPU compute has historically been sold and deployed. It's more like buying an appliance than buying components, and it gives NVIDIA significantly more control over the full-stack performance story. When NVIDIA says Vera Rubin delivers the lowest token cost, they're talking about a system that's been tuned end-to-end, not a GPU dropped into a third-party server.

Partners currently running NVL72 hardware include CoreWeave, which has become one of the most important AI cloud providers in the market, and Google, which is both a cloud competitor and a major AI infrastructure customer simultaneously. That Google is buying Vera Rubin racks while also developing its own TPU line says something about how dominant NVIDIA's position remains, even as alternatives multiply.

The future, in 3 minutes a day. The biggest tech story explained every morning, free. Get the briefing →

The Performance Per Watt Story

Power consumption is increasingly the binding constraint on AI infrastructure expansion. Data centres are running into electricity supply limits, and hyperscalers are signing long-term energy deals and even investing in nuclear power to keep up with demand. In that context, performance per watt isn't just an efficiency metric, it's the thing that determines whether you can actually build more AI capacity without running out of power.

NVIDIA has been consistent in positioning Vera Rubin as a generational leap on this front compared to the Hopper generation. The Blackwell architecture (which sits between Hopper and Rubin in the roadmap) already brought significant efficiency gains, and Rubin is meant to push further. Whether those claims hold up in real-world deployments at CoreWeave and Google will become clearer over the coming months as operators publish their own benchmarks.

What This Means for the Market

For companies building AI products, the arrival of Vera Rubin in production is good news in principle. More efficient hardware at the top end should eventually flow through to lower inference costs across the cloud providers, which should make AI-powered products cheaper to run. The lag between new hardware launching and those savings appearing in cloud pricing is typically six to twelve months.

For NVIDIA as a business, Vera Rubin represents continued dominance of the AI training and inference market despite AMD, Intel, and various hyperscaler custom silicon efforts trying to carve out share. The fact that Google is a Rubin customer even while running its own TPU fleet is probably the most telling data point of all.

The race for compute efficiency is not slowing down. Vera Rubin is the current high-water mark, but there are already hints of what comes next on NVIDIA's roadmap. For now though, the NVL72 is what the AI industry runs on.

Get the briefing, free

The biggest tech story, explained in 3 minutes every weekday. Choose your briefings →

Free. No spam. Unsubscribe in one click.