AMD Helios vs Nvidia Vera Rubin: The Memory Number Decides This One
Key takeaways
- Helios packs 72 Instinct MI455X accelerators and sixth-generation Epyc 9006 CPUs into one rack with 31TB of HBM4
- Nvidia's DGX Vera Rubin NVL72 counters with 20.7TB of GPU memory and up to 3,600 PFLOPS of NVFP4 inference
- AMD trails on headline inference FLOPS and leads on memory capacity by roughly 50 percent, which is a deliberate bet on memory-bound serving
- Helios runs on OCP Open Rack Wide, UALink over Ethernet and standard Broadcom switching, against Nvidia's vertically integrated stack
31TB against 20.7TB. That is the number AMD wants read first, and it is why AMD Helios vs Nvidia Vera Rubin is a more interesting fight than the last four years of chip announcements. AMD is not selling a competitor to a GPU. It is selling a competitor to a rack.
What is actually in the box
Helios pairs 72 Instinct MI455X accelerators with sixth-generation Epyc 9006 CPUs, Pensando networking and the ROCm software stack, sold as one integrated system. Per rack that works out to 31TB of HBM4, somewhere between 1.4 and 1.67 PB/s of aggregate memory bandwidth, and around 2.9 exaFLOPS at MXFP4. Power draw lands in the region of 225 to 245kW.
Nvidia's DGX Vera Rubin NVL72 is the target. It carries 72 Rubin GPUs, 36 Vera CPUs, sixth-generation NVLink, 20.7TB of total GPU memory and up to 3,600 PFLOPS of NVFP4 inference.
On raw inference throughput Nvidia is ahead. On memory AMD is ahead by roughly half again. AMD picked that trade deliberately, and the reasoning is the interesting part.
Why memory is the number that matters
A serving cluster that runs out of memory before it runs out of compute has idle silicon in it. Large models with long context windows spend most of their time moving weights and key-value caches around rather than doing arithmetic. In plain terms, the accelerator is starving rather than thinking.
AMD has read that constraint and built to it. More HBM per rack means larger models fit without splitting across nodes, and longer contexts fit without eviction. The compute deficit only bites on workloads that were compute-bound to begin with, and fewer of those exist in production inference than the FLOPS marketing suggests. Anyone tracking how much power AI data centres actually use will recognise the pattern, because the headline number rarely describes the bottleneck.
The part worth sitting with
The second story here is architectural politics. AMD built Helios on OCP Open Rack Wide, UALink running over Ethernet, Ultra Ethernet Consortium networking and standard Broadcom switching. Nvidia's stack is integrated end to end, which is why it performs the way it does and also why procurement teams have spent two years worrying about it.
Every buyer who has been asked to justify single-vendor dependency now has a spec sheet to point at. That is a commercial argument rather than a technical one, and it may carry more weight than the memory figure. AMD has been chipping away at Nvidia from both directions, having already taken close to 30 percent of x86 market share this year, and the same open-standards question runs through the wider consolidation story, including Nvidia's acquisition of Hugging Face.
What to watch
Neither system has independent benchmarks attached yet. Vendor figures for memory bandwidth and low-precision throughput get measured under conditions the vendor picked, and MXFP4 against NVFP4 is not a like-for-like comparison of formats. The number to wait for is tokens per second per watt, on a model somebody actually serves, at a context length somebody actually uses. Until that exists this is a fight between two spec sheets, and the spec sheets disagree about what the job is.