Future TechnologyFuture Technology
← Back to archive
Newsletter

Nvidia won training. Inference is a different problem entirely

10 September 2026
Nvidia won training. Inference is a different problem entirely
Future Technology

New article published

COMPUTING

Nvidia won training. Inference is a different problem entirely

Training a frontier model is a one off cost. Inference is a bill that arrives every time somebody uses it. That single difference explains why hardware announcements now lead with memory numbers instead of teraflops.

Key Takeaways

  • Training is a one off cost per model generation, while inference is paid on every single request
  • Training rewards raw compute throughput, but the decode phase of inference is bottlenecked by memory bandwidth and capacity
  • AMD Helios ships 72 accelerators with 31 terabytes of memory against 20.7 terabytes in the Nvidia equivalent, and loses on headline performance
  • If a hardware launch leads with FLOPS it is a training pitch, and if it leads with terabytes it is aimed at inference

You received this because you subscribe to Future Technology.

Unsubscribe