FTFuture Technology
HARDWARE

NVIDIA's Vera Rubin Is Already at Scale, and the Cost Numbers Are Striking

· 3 min read · By Nath Connell

Key takeaways

  • Vera Rubin NVL72 racks are in live production deployment at CoreWeave and Google
  • The NVL72 configuration uses 72 Rubin GPUs interconnected via NVLink in a single rack
  • NVIDIA claims Vera Rubin delivers the lowest token cost of any competing platform currently available

NVIDIA's Vera Rubin platform is ramping faster than anyone in the industry expected, and the performance numbers coming out of early deployments are worth paying attention to. Vera Rubin NVL72 racks are now running at partners including CoreWeave and Google, with production ramping across multiple hyperscale customers. The headline claim from NVIDIA is that Vera Rubin delivers the lowest token cost among its competitors, a metric that, if it holds up, has significant implications for how AI services are priced and who can afford to run frontier models.

Token cost, for those not steeped in AI infrastructure economics, is exactly what it sounds like: how much it costs to process a given amount of text or other data through a large language model. It is the fundamental unit of economics for AI services, and it is what determines whether running a frontier AI model is viable for a startup, a hospital, or a government department, or whether it remains the exclusive domain of organisations with deep pockets.

What NVL72 Actually Is

The Vera Rubin NVL72 is NVIDIA's current flagship server configuration. The NVL72 designation indicates a rack configuration with 72 Rubin GPUs, interconnected with NVIDIA's NVLink technology for high-bandwidth communication between chips. This tight integration means that models which do not fit within the memory of a single GPU can be spread across multiple chips without the severe performance penalties that arise when chips have to communicate over slower interfaces.

This architecture is specifically designed for the massive models that frontier AI now requires. Earlier GPU clusters could handle large models, but they required complex and expensive network infrastructure to move data between chips quickly enough. NVLink integrated directly into the Vera Rubin rack configuration shifts that bottleneck significantly, which is a large part of where the efficiency gains come from.

The Vera CPU component of the platform is also worth noting. NVIDIA's custom ARM-based Vera CPU is designed to work in close coordination with the Rubin GPUs, handling the parts of AI workloads that do not benefit from GPU parallelism without creating the bottlenecks that arise when pairing GPU clusters with general-purpose server CPUs from AMD or Intel.

The future, in 3 minutes a day. The biggest tech story explained every morning, free. Get the briefing →

CoreWeave and Google as Reference Deployments

The fact that both CoreWeave and Google are named as live deployment partners matters beyond just validating the technology. CoreWeave is the AI-focused cloud provider that has become one of the most important independent GPU cloud operators in the world, with a business model built almost entirely around providing AI compute capacity. When CoreWeave chooses a platform, it is an explicit performance and economics judgement, because their entire business depends on running AI workloads as efficiently as possible.

Google is a different kind of signal. Google designs its own AI accelerators, the Tensor Processing Units, and has been doing so since 2015. The fact that Google is also deploying Vera Rubin suggests that even companies with strong in-house silicon programmes see value in NVIDIA's platform for certain workloads, rather than relying exclusively on their own chips.

What Lowest Token Cost Actually Means

NVIDIA's claim that Vera Rubin achieves the lowest token cost needs to be understood in context. Token cost is a function of multiple variables: raw compute throughput, memory bandwidth, power consumption, and the efficiency of the software stack running on the hardware. NVIDIA's claim is that the combination of Rubin GPU architecture, NVLink interconnect, and NVIDIA's software optimisations produces a better outcome on this composite metric than alternatives.

If the claim holds, the downstream effect on AI service pricing could be substantial. The cost of running AI inference has been falling rapidly over the past three years, driven by a combination of hardware improvements and software optimisation. Cheaper inference means more use cases become economically viable, and it shifts competitive dynamics in the AI services market towards whoever can pass those cost reductions on to customers most effectively.

The ramp is still early, and independent benchmarks at scale will matter more than launch claims. But the deployment momentum, with two major hyperscale partners already running production workloads, suggests that Vera Rubin is not just a paper announcement.

Get the briefing, free

The biggest tech story, explained in 3 minutes every weekday. Choose your briefings →

Free. No spam. Unsubscribe in one click.