Free tools · AI
LLM VRAM calculator: will this model run on my GPU or Mac?
An 8B model at Q4_K_M needs about 6 GB and runs at well over 100 tokens a second on a 16 GB card. A 70B model needs about 48 GB, which is two GPUs or a 64 GB Mac. Pick your model and hardware below for the exact numbers, a speed estimate and the cheapest UK hardware that fits.
- Weights
- Context cache
- Overhead
- Total needed
- Available
- Speed estimate
- Largest context that fits
Cheapest GPU that fits:
Cheapest Mac or unified-memory machine that fits:
Hardware links may earn us a commission at no cost to you. Prices are UK bands at the verified date, not live quotes. See our affiliate disclosure.
Model and GPU prices move monthly. One email when we refresh this table, plus the daily briefing.
How it works
Three things have to fit in memory at once: the model weights, the context cache (the key and value tensors for every token in the conversation) and the runtime's own overhead. Weights are the parameter count times the bits per weight of the quantisation, divided by eight. The context cache is 2 x layers x KV heads x head dimension x context length x bytes per element. Overhead is taken as the larger of 1 GB or 10 percent of the weights, plus half a gigabyte. The verdict is "fits" when the total is at or below 92 percent of available memory.
Speed is bounded by memory bandwidth. Every generated token reads the active weights once, so the ceiling is bandwidth divided by bytes per token, and real runtimes land at 60 to 70 percent of that. Mixture-of-experts models only read their active experts, which is why a 109B Llama 4 Scout can be quicker than a 70B dense model despite needing more memory.
Assumptions and data
Every default in the calculator comes from the table below, verified on . Model shapes are read from each model's published configuration, hardware memory and bandwidth from the makers' specification pages, and prices are indicative UK bands.
| Item | Value | Source |
|---|
Worked example
Qwen3 32B at Q4_K_M has 32.8 billion parameters at 4.85 bits each, so the weights take about 19.9 GB. With an 8k context and FP16 cache, the context cache is 2 x 64 layers x 8 KV heads x 128 x 8,192 x 2 bytes, about 2.1 GB. Overhead adds 2.5 GB, for a total near 24.5 GB. That is over a 24 GB RTX 4090, so the verdict is "tight" at best; drop the context to 4k or the cache to Q8 and it fits. On a Mac Studio with 36 GB, 27 GB is usable at the default share, and it fits with room to spare at roughly 10 to 12 tokens a second.
Frequently asked questions
How much VRAM do I need to run a 70B model?
At Q4_K_M, about 43 GB for the weights alone, plus a few gigabytes for context and overhead. That means two 24 GB GPUs, a 32 GB RTX 5090 with layers offloaded, or a Mac with 64 GB or more of unified memory.
Can I run a local LLM on 8 GB of VRAM?
Yes. Models up to about 8B parameters at Q4_K_M use roughly 5 to 6 GB with an 8k context. 12B to 14B models fit at Q3 or with a short context. Anything larger needs offloading to system RAM, which works but is slow.
Is a Mac or a GPU better for local models?
A GPU has far more memory bandwidth, so it is faster for anything that fits in its VRAM. A Mac with 64 GB or more can load models that no single consumer GPU can, but runs them more slowly. The result card shows the cheapest of each that fits your choice.
What does quantisation do to quality?
Q4_K_M roughly quarters the memory of a 16-bit model with a small, usually acceptable loss. Q3 and Q2 save more but the loss becomes noticeable, especially on small models. Q8 is close to lossless and worth it if the memory is there.
Why is my real memory use different?
Runtimes differ: llama.cpp, Ollama, vLLM and MLX allocate context and buffers in their own ways, and drivers and other apps take memory too. Treat the estimate as plus or minus 15 percent, and the speed as single-stream decoding without prompt processing.
Related reading
Changelog
- 11 October 2026: first release, 21 models, 17 hardware presets.