NVIDIA Vera Rubin Is Ramping Fast and Already Driving Down the Cost of AI Tokens
Key takeaways
- NVIDIA Vera Rubin NVL72 racks are in production at CoreWeave and Google as of mid-2026
- The NVL72 configuration packs 72 Rubin GPUs connected via NVLink fabric in a single rack
- NVIDIA claims Vera Rubin delivers the lowest token cost and best performance per watt of any current platform
- Token cost per million tokens is the primary economic metric for enterprises running AI inference at scale
NVIDIA's Vera Rubin architecture has moved from announcement to meaningful production scale, and the early performance story is compelling. Vera Rubin NVL72 racks are now running at partners including CoreWeave and Google, and NVIDIA is making a specific claim that matters a great deal to the economics of AI: Vera Rubin is delivering the lowest token cost of any platform currently available for large-scale inference.
Token cost is the metric that increasingly defines AI economics. Every time a large language model generates a word, an image, or a piece of code, it is consuming compute measured in tokens. For AI companies and enterprises running inference at scale, the cost per million tokens is what determines whether a given AI application is financially viable. Driving that cost down is not just a benchmark achievement. It is what makes a much wider range of AI applications economically practical.
What the NVL72 Configuration Brings
The NVL72 rack packs 72 Rubin GPUs into a single system, connected via NVIDIA's NVLink fabric. This dense interconnect is key to the performance claim. When dozens of GPUs can communicate with each other at extremely high bandwidth and low latency, large AI models can be split across the hardware more efficiently, reducing the idle time that eats into both performance and cost.
Vera Rubin also introduces improvements in performance per watt, which matters enormously as AI data centres push up against power infrastructure limits. The energy cost of running AI at scale has become one of the defining constraints on the industry, and a chip that delivers more computation per watt is directly valuable even before you consider raw performance. If Vera Rubin is genuinely more efficient per watt than its predecessors, hyperscalers have a strong incentive to refresh their hardware faster than typical cycles would suggest.
The Competitive Picture
NVIDIA's position in AI inference hardware has been challenged more seriously in the past 18 months than at any point in the company's recent history. AMD has been gaining ground with its MI-series chips, Google's TPU v5 has proven highly cost-effective for its own workloads, and a wave of AI chip startups including Groq, Cerebras, and others have attracted serious investment and customer interest by targeting the inference market specifically.
Against that backdrop, the Vera Rubin performance-per-watt and token cost claims are important because they are the specific metrics where alternatives have been making inroads. If NVIDIA can demonstrate that Vera Rubin outperforms or matches competitors on token cost while benefiting from the broader CUDA software ecosystem and the trust established by years of deployment at hyperscaler scale, the competitive challenge becomes significantly harder to press.
CoreWeave and Google running production NVL72 workloads is meaningful validation. These are not reference customers who receive preferential treatment and report inflated numbers. Both companies have strong commercial incentives to deploy whatever hardware delivers the best economics for their customers, and both have the engineering depth to benchmark alternatives rigorously.
What Comes Next
NVIDIA has historically maintained a roughly two-year cadence between major architecture generations. Vera Rubin's production ramp in mid-2026 suggests the next generation is likely to be announced in late 2027 or early 2028. In the meantime, the focus will shift to software: improving inference efficiency through better batching, quantisation, and speculative decoding techniques that can extract more performance from existing hardware.
For organisations making AI infrastructure decisions now, Vera Rubin represents a mature, validated option with a clear cost story. The question is whether the performance advantages justify the capital expenditure over alternatives that have become increasingly capable. Based on what we know about production deployments to date, NVIDIA's answer to that question appears to be yes, and the partners running NVL72 systems at scale are the most credible witnesses to that claim.