An 8B model at Q4_K_M needs about 6 GB and runs at well over 100 tokens a second on a 16 GB card. A 70B model needs about 48 GB, which is two GPUs or a 64 GB Mac. Pick your model and hardware below for the exact numbers, a speed estimate and the cheapest UK hardware that fits.

How it works

Three things have to fit in memory at once: the model weights, the context cache (the key and value tensors for every token in the conversation) and the runtime's own overhead. Weights are the parameter count times the bits per weight of the quantisation, divided by eight. The context cache is 2 x layers x KV heads x head dimension x context length x bytes per element. Overhead is taken as the larger of 1 GB or 10 percent of the weights, plus half a gigabyte. The verdict is "fits" when the total is at or below 92 percent of available memory.

Speed is bounded by memory bandwidth. Every generated token reads the active weights once, so the ceiling is bandwidth divided by bytes per token, and real runtimes land at 60 to 70 percent of that. Mixture-of-experts models only read their active experts, which is why a 109B Llama 4 Scout can be quicker than a 70B dense model despite needing more memory.

Assumptions and data

Every default in the calculator comes from the table below, verified on . Model shapes are read from each model's published configuration, hardware memory and bandwidth from the makers' specification pages, and prices are indicative UK bands.

ItemValueSource

Worked example

Qwen3 32B at Q4_K_M has 32.8 billion parameters at 4.85 bits each, so the weights take about 19.9 GB. With an 8k context and FP16 cache, the context cache is 2 x 64 layers x 8 KV heads x 128 x 8,192 x 2 bytes, about 2.1 GB. Overhead adds 2.5 GB, for a total near 24.5 GB. That is over a 24 GB RTX 4090, so the verdict is "tight" at best; drop the context to 4k or the cache to Q8 and it fits. On a Mac Studio with 36 GB, 27 GB is usable at the default share, and it fits with room to spare at roughly 10 to 12 tokens a second.

Frequently asked questions

How much VRAM do I need to run a 70B model?

At Q4_K_M, about 43 GB for the weights alone, plus a few gigabytes for context and overhead. That means two 24 GB GPUs, a 32 GB RTX 5090 with layers offloaded, or a Mac with 64 GB or more of unified memory.

Can I run a local LLM on 8 GB of VRAM?

Yes. Models up to about 8B parameters at Q4_K_M use roughly 5 to 6 GB with an 8k context. 12B to 14B models fit at Q3 or with a short context. Anything larger needs offloading to system RAM, which works but is slow.

Is a Mac or a GPU better for local models?

A GPU has far more memory bandwidth, so it is faster for anything that fits in its VRAM. A Mac with 64 GB or more can load models that no single consumer GPU can, but runs them more slowly. The result card shows the cheapest of each that fits your choice.

What does quantisation do to quality?

Q4_K_M roughly quarters the memory of a 16-bit model with a small, usually acceptable loss. Q3 and Q2 save more but the loss becomes noticeable, especially on small models. Q8 is close to lossless and worth it if the memory is there.

Why is my real memory use different?

Runtimes differ: llama.cpp, Ollama, vLLM and MLX allocate context and buffers in their own ways, and drivers and other apps take memory too. Treat the estimate as plus or minus 15 percent, and the speed as single-stream decoding without prompt processing.

Related reading

Changelog

  • 11 October 2026: first release, 21 models, 17 hardware presets.