Computing

NPU vs GPU vs CPU: What actually runs your AI

(1 month ago) · 7 min read · By Future Technology · Edited by Nath Connell

Key takeaways

  • A CPU handles logic, a GPU handles flexible parallel maths, an NPU handles one narrow shape of maths very efficiently
  • A TOPS figure is a theoretical peak at low precision and cannot be compared across vendors
  • Memory bandwidth, not TOPS, sets how many tokens per second you actually get

Every laptop spec sheet sold in 2026 carries a TOPS figure, and almost nobody selling you the laptop can explain what it measures. The NPU vs GPU vs CPU question sits underneath every one of those purchases, and the answer is less about which is fastest than about which is doing the work.

What is the difference between a CPU, a GPU and an NPU?

All three are processors, but they are built around different bets about what work will arrive. A CPU (central processing unit) bets on variety. A GPU (graphics processing unit) bets on repetition. An NPU (neural processing unit) bets on a single kind of repetition, and optimises everything else away.

Knowing which bet each chip made tells you most of what you need to know when a vendor says its laptop has "AI acceleration". The table below sets the three side by side.

ChipCore designBest atWeak atWhere you find it
CPUA handful of powerful, flexible coresBranching logic, sequential work, anything unpredictableLarge amounts of parallel mathsEvery computer and phone
GPUThousands of simple coresMatrix multiplication and other parallel mathsPower efficiencyGaming PCs, workstations, servers, laptops
NPUFixed-function siliconLow-precision matrix maths at very low powerAnything outside that narrow jobLaptops and phones

What is a CPU good at?

A CPU has a handful of powerful, flexible cores. It is brilliant at branching logic, sequential work and anything unpredictable. Think of the operating system deciding which application gets time, a spreadsheet recalculating, or a browser parsing a page: each step depends on the one before it, so you cannot simply throw more cores at it.

A CPU will run an AI model, but slowly, and it still has to run everything else on the machine while it does. That is why a model running on the CPU alone tends to make the whole computer feel sluggish.

What is a GPU good at?

A GPU has thousands of simple cores doing the same operation on different data. That is exactly the shape of a matrix multiplication, which is exactly what a neural network is made of. Rendering a game frame and running a language model turn out to be close cousins, which is why graphics cards ended up powering the whole AI boom.

The catch is that a GPU is flexible, fast and thirsty. Every one of those thousands of cores draws power, so a GPU doing sustained AI work in a thin laptop drains the battery quickly and produces heat. For a deeper look at how the same trade-off plays out in data centres, see our piece on AI inference chips versus training chips.

What is an NPU good at?

An NPU is fixed-function silicon built for one narrow job: low-precision matrix maths, at very low power. It is faster per watt than a GPU on the work it was designed for and awkward at anything outside that. Because the circuitry is dedicated rather than general, nothing is wasted on features the job never uses.

That power figure is why NPUs appear in laptops and phones rather than servers. A phone has a battery measured in hours, not a mains socket, so a chip that can recognise a face or transcribe speech without draining the battery is worth the silicon. In a data centre, where power is plentiful and flexibility matters more, the GPU keeps its place.

What does a TOPS number actually tell you?

TOPS means trillions of operations per second. It is almost always quoted at INT8 or lower precision, and it is almost always a theoretical peak rather than anything you will sustain. INT8 means each number in the calculation is stored in 8 bits, which is coarse but good enough for many models and much cheaper to compute than full-precision numbers.

The figure does not tell you the precision your model needs, whether the number counts the NPU alone or the whole chip, or whether your software can reach the hardware at all. A large TOPS figure is no use if the app you want to run was written for the GPU and never calls the NPU.

Can you compare TOPS across vendors?

Comparing TOPS across vendors is close to meaningless. Two chips with the same number can differ by a wide margin on the same workload, which is worth remembering when reading the Snapdragon C machines or Intel Panther Lake parts against each other.

The reasons are practical. Vendors count different units, quote different precisions and rely on different software stacks. If you are weighing laptops on this basis, our explainer on Arm laptops going mainstream covers why the chip architecture underneath matters as much as the headline number.

Why does memory bandwidth matter more than TOPS?

Generating a single token requires streaming the model weights through the compute unit. A token is a small chunk of text, roughly a short word or part of one, and a model produces its answer one token at a time.

An 8GB model means 8GB moved, per token, every token. That makes memory bandwidth the real ceiling on local inference, not raw compute. A chip can have enormous arithmetic capacity and still sit idle, waiting for the next slice of weights to arrive from memory.

The practical consequence is that two machines with identical TOPS can generate text at very different speeds if one has faster memory. Our guide to HBM, GDDR and DDR memory explains why different memory types deliver such different bandwidth, and it is more useful for predicting speed than any TOPS figure. If you want to put this into practice, how to run a local LLM on a Mac in 2026 walks through the setup.

Why does the Apple M6 split make the point?

Apple gave the entry M6 a dual 16-core Neural Engine and the more expensive M5 Pro a single one, according to Apple Newsroom and MacRumors. That reads backwards until you notice who buys each.

Pro buyers grade video and compile code, which leans on CPU and GPU cores. The entry Mac mini is being sold as a local AI box, so it got the NPU. Apple has matched the silicon to the job each buyer is likely to do, rather than simply giving the dearest chip the biggest of everything.

How should you read a spec sheet?

Work out what you are asking the machine to do, then look at the unit that does it. If your work is editing video or compiling software, the CPU and GPU core counts matter. If your work is running small models locally, look at the NPU, and then look at the memory bandwidth feeding it.

The headline number is rarely the one that matters. A sensible approach is to ask three questions of any AI-branded machine: which unit does the work I care about, can my software reach it, and how quickly can memory feed it. A seller who can answer all three is rarer than one who can recite a TOPS figure.

Key takeaways

  • A CPU handles branching logic and sequential work, a GPU handles flexible parallel maths with thousands of simple cores, and an NPU handles one narrow shape of low-precision maths very efficiently.
  • An NPU is faster per watt than a GPU on the work it was designed for, which is why it appears in laptops and phones rather than servers.
  • A TOPS figure is a theoretical peak, usually quoted at INT8 or lower precision, and cannot be sensibly compared across vendors.
  • Memory bandwidth, not TOPS, sets the ceiling on local inference, because an 8GB model means 8GB moved for every token.
  • Apple gave the entry M6 a dual 16-core Neural Engine and the M5 Pro a single one, because the entry Mac mini is sold as a local AI box.

More from Future Technology