Nvidia won training. Inference is a different problem entirely
Future Technology
New article published
COMPUTING
Nvidia won training. Inference is a different problem entirely
Training a frontier model is a one off cost. Inference is a bill that arrives every time somebody uses it. That single difference explains why hardware announcements now lead with memory numbers instead of teraflops.
Key Takeaways
- Training is a one off cost per model generation, while inference is paid on every single request
- Training rewards raw compute throughput, but the decode phase of inference is bottlenecked by memory bandwidth and capacity
- AMD Helios ships 72 accelerators with 31 terabytes of memory against 20.7 terabytes in the Nvidia equivalent, and loses on headline performance
- If a hardware launch leads with FLOPS it is a training pitch, and if it leads with terabytes it is aimed at inference
You received this because you subscribe to Future Technology.