Hardware

Nvidia GPUs, Google TPUs and AWS Trainium all run AI, but only one is rentable

(5 days ago) · 4 min read · By Future Technology

Key takeaways

  • Nvidia GPUs are the only chip on this list you can rent as bare silicon
  • Google's seventh generation Ironwood TPU is roughly six years more mature than Trainium
  • AWS has over 500,000 Trainium2 chips in production and killed the separate Inferentia line
  • Custom ASICs are growing around 44.6 percent a year and could be 45 percent of AI chips by 2028

Four chip families run almost all of the AI workload in 2026. You can rent exactly one of them. That constraint explains more about the market than any benchmark does.

GPU vs TPU vs Trainium, in plain terms

Nvidia GPU

General purpose, available to anyone with a credit card, and still the default for training. Every major cloud rents them and every framework targets them, which is why any company that is not a hyperscaler ends up building on them. Nvidia sells access as much as it sells silicon.

Google TPU

Seventh generation Ironwood is in production. It is roughly six years more mature than Trainium and runs at around ten times the cluster scale. The eighth generation, announced at Google Cloud Next in April 2026, is the first time Google has fielded two architectures in one generation, one tuned for training and one for inference. It is built on TSMC N3 and hosted on Google's own Axion Arm CPUs.

AWS Trainium

Over 500,000 Trainium2 chips are in production, which is the largest hyperscaler ASIC fleet by unit count. Trainium3 ships at 2.52 PFLOPS FP8 per chip with 144GB of HBM3e. AWS retired the separate Inferentia line, because inference workloads have started to look computationally like training.

Everyone else's ASIC

Microsoft Maia, Meta MTIA, and OpenAI working with Broadcom. The custom ASIC market is growing at roughly 44.6 percent a year and could account for 45 percent of AI chips by 2028.

The catch nobody puts in the comparison table

None of the four hyperscaler chips are rentable as bare silicon. They are captive to the companies that built them. You can rent a Google Cloud instance that happens to sit on a TPU, but you cannot buy a TPU, put it in your rack, or move your workload to a different provider running the same part.

So the practical comparison for anyone outside AWS, Google, Microsoft and Meta is not four way. It is Nvidia, and then a set of chips that quietly set the cost floor your Nvidia bill is measured against.

Why everything is drifting toward inference

Inference is now roughly two thirds of AI compute spend. That is the number that explains the TPU split into training and inference variants, the retirement of Inferentia, and Qualcomm's move into AWS data centres with inference silicon.

A purpose built inference chip can beat a general purpose GPU on cost per token by a wide margin, because inference is a narrower problem with predictable shapes. Training rewards flexibility. Inference rewards specialisation, and specialisation is what an ASIC is for.

What this means if you run models locally

Very little of the above is buyable, so the consumer question stays a VRAM question. Our local LLM VRAM tier list is the practical version of this comparison for anyone with one machine rather than one data centre.

The gap is widening in the other direction too. Nvidia is sitting on 279 billion dollars in purchase commitments through 2027, which is a supply position, not a demand signal.

What to watch

Whether any hyperscaler offers its ASIC as a rentable, portable instance type rather than a managed service. That is the moment the comparison table becomes useful to people who do not work at one of four companies. Nothing announced so far suggests it is coming.

More from Future Technology