NVIDIA Vera Rubin Is Here and Already Running at Scale With CoreWeave and Google
Key takeaways
- Vera Rubin NVL72 racks are in production at partners including CoreWeave and Google as of July 2026
- NVL72 connects 72 GPUs as a single logical unit via NVLink, optimised for large model inference
- NVIDIA claims Vera Rubin delivers the lowest token cost of any hardware it has shipped to date
- Performance per watt improvements are being validated in live production deployments
NVIDIA's Vera Rubin architecture has moved from announcement to production, and the ramp is happening fast. As of late July 2026, Vera Rubin NVL72 racks are already running at partners including CoreWeave and Google, with gigascale deployment underway. For a chip generation that was still being talked about in future-tense just months ago, that is a strikingly quick transition to real-world infrastructure.
Vera Rubin is NVIDIA's successor to the Hopper and Blackwell generations, and it represents a meaningful step forward on two metrics that data centre operators care about most right now: performance per watt and token cost. As AI inference workloads scale, the economics of running large language models become increasingly dependent on how cheaply you can generate a token. NVIDIA is positioning Vera Rubin as the architecture that drives that cost lower than anything before it.
What Makes Vera Rubin Different
The NVL72 form factor is significant. NVL stands for NVLink-connected, meaning 72 GPUs operate as a single logical unit rather than communicating over slower interconnects. This kind of scale-up architecture is particularly well suited to large model inference, where the model is too big to fit on a single GPU and the communication overhead between chips becomes a bottleneck.
Vera Rubin also introduces updated memory configurations designed to work with the next generation of HBM, keeping data flowing fast enough to match the raw compute throughput of the GPU array. The combination of high interconnect bandwidth and updated memory is what allows NVIDIA to credibly claim the performance per watt improvements partners are starting to validate in production.
The choice of CoreWeave and Google as the first visible production partners is telling. CoreWeave has built its entire business around being the fastest to deploy new NVIDIA hardware at scale, making it a natural first mover. Google, operating at hyperscaler scale, validates that Vera Rubin can handle the kind of sustained, mixed workload profile that enterprise customers eventually see reflected in their cloud costs.
The Economics of Tokens
Token cost is becoming the primary metric by which AI infrastructure is evaluated, and it is worth understanding why. When a business runs an AI application, whether that is a customer service agent, a code completion tool, or a document processing pipeline, the marginal cost of that application is dominated by the cost of generating tokens. Cheaper tokens mean lower operating costs for AI applications, which means more applications become economically viable.
NVIDIA claiming Vera Rubin delivers the lowest token cost for partners worldwide is not just a performance brag. It is a direct argument that building on Vera Rubin makes AI applications more profitable. That argument lands differently now than it would have in 2023, when most enterprises were still running proof-of-concept projects. In 2026, the cost of AI inference is a real line item on real budgets.
What This Means for the Competitive Landscape
AMD, Intel, and a range of custom silicon makers including Google's own TPU team are all competing in this space. But the speed at which NVIDIA moves from announcement to gigascale production continues to be one of the company's most durable competitive advantages. When Vera Rubin NVL72 racks are already running at partners before most competitors have shipped meaningful volume of their current-generation alternatives, it creates a window in which NVIDIA's architecture is the only realistic choice for teams that cannot wait.
Software lock-in through CUDA remains real, but the more immediate moat right now is production velocity. NVIDIA's supply chain, manufacturing partnerships, and partner ecosystem allow it to ramp new hardware faster than the competition can match, and every month of lead time translates directly into data centre contracts signed on NVIDIA terms.
For developers and infrastructure teams, the practical question is when Vera Rubin capacity becomes available through major cloud providers and what the pricing looks like versus Blackwell. Based on NVIDIA's track record, that access will come sooner than expected. The gigascale ramp has already started.