How to run a local LLM on a Mac in 2026
Key takeaways
- Unified memory is the ceiling, not the GPU. Parameters in billions times 0.6 gives roughly the gigabytes needed at 4-bit
- Q4_K_M quantisation is the usual sweet spot between output quality and memory use
- Anything needing current information or a frontier-scale model still belongs in the cloud
A Mac with 32GB of unified memory will run a 20 billion parameter model at 4-bit without complaint. A 16GB Mac will not. That one number decides almost everything else about running a local LLM (large language model) on a Mac, so start there rather than with the model you want.
This guide covers the four decisions in order: how much memory you have, which tool runs the model, how heavily the model is compressed, and which size suits the job. It also covers when a local model is the wrong tool, and what to buy if you are starting from scratch.
How much memory do you need to run a local LLM on a Mac?
On Apple silicon the GPU is rarely the limit. Unified memory is. Unified memory means the CPU and GPU share one pool, so the model's weights sit in the same place whichever chip is working on them. There is no separate, smaller pool of graphics memory to run out of first.
The rough working rule: take the model size in billions of parameters, multiply by 0.6, and that is roughly the gigabytes of memory the weights need at 4-bit quantisation. A 14B model wants about 8.4GB before anything else loads. The same rule puts a 20B model at around 12GB, which is why it sits happily on a 32GB machine with room left for macOS and your other apps.
Then leave headroom. Context windows are stored on top of the weights, and a long one is not free. The context window is the amount of text the model can hold in mind at once, covering your prompt, the conversation so far and its own reply. Longer contexts take more memory, so a model that just fits with a short chat can stall once you paste in a long document.
If you have already read our VRAM requirements breakdown, the same arithmetic applies here with unified memory in place of dedicated VRAM. For a tier-by-tier view, see what you can actually run locally by how much RAM you have.
Which runner should you use: Ollama, LM Studio or llama.cpp?
Three options cover almost everyone. Ollama is the fastest way in: one command pulls a model and starts talking to it. LM Studio gives you a GUI, a model browser and a local API server. llama.cpp is what the other two are built on, and it is the right choice if you want every knob exposed.
| Runner | Best for | Interface | Main strength |
|---|---|---|---|
| Ollama | Getting started quickly | Command line | One command pulls and runs a model |
| LM Studio | People who prefer windows to terminals | GUI | Model browser and local API server |
| llama.cpp | Tinkerers who want full control | Command line | Every setting exposed |
All three are free to download from their own sites: Ollama, LM Studio and llama.cpp. If you later want to spread work across several machines, our piece on vLLM, Ollama and other local inference tools explains where each fits.
Which quantisation should you choose?
Quantisation shrinks a model by storing each weight with fewer bits. It is the reason a model that would otherwise need a server can fit on a laptop.
Q4_K_M is the default answer. It is the usual sweet spot between output quality and memory use, and it is the setting the 0.6 rule above assumes.
Moving up to 8-bit roughly doubles the memory for a quality difference most people cannot reliably pick out in ordinary use. On the 14B example, that takes you from about 8.4GB to somewhere near 17GB, which is a lot to pay for a gain you may never notice.
Dropping to 3-bit saves memory and starts to show, particularly on longer reasoning chains. If a model only fits at 3-bit, it is usually better to choose a smaller model at 4-bit instead.
Which model size suits which job?
Small models in the 3B to 8B range are genuinely good at autocomplete, summarising, reformatting and tidying text, and they run fast enough to feel instant. They are also forgiving on memory, so they work on almost any recent Mac.
Mid-size models around 14B to 32B handle reasoning tasks passably. Expect them to manage drafting, code explanation and multi-step questions, but not to match the best hosted models on hard problems. They are also where memory starts to matter, since a 32B model at 4-bit needs roughly 19GB for the weights alone.
The frontier stays in the cloud, and pretending otherwise wastes an afternoon. If you want a sense of how far a single desktop machine can stretch, our guide to running a 30B agent locally sets out what that takes.
When does a local model lose to the cloud?
A local model has no idea what happened this morning, cannot search, and will not hold a million-token context. For current information, very long documents or best-in-class reasoning, a hosted model still wins.
Local wins on privacy, on latency, and on the fact that there is no monthly bill. Your prompts never leave the machine, replies start without a network round trip, and you pay nothing per request once the hardware is bought.
This also affects how much access you hand to automated tools. Our comparison of a local AI agent and a cloud agent looks at which should be allowed near your accounts.
What should you buy if you are starting from scratch?
The new Mac mini M6 ships on 22 September 2026 and is being pitched squarely at this use case. Our Mac mini M6 versus Mac Studio comparison covers which one suits local models.
If you would rather not wait, the outgoing M4 model is available now in a 24GB configuration, which is a comfortable place to sit. On the Windows side the RTX Spark route solves the same problem differently and costs more.
Whichever you choose, buy the memory rather than the cores. Cores make a model faster. Memory decides whether it runs at all. A cheaper chip with more memory will usually serve you better than a faster chip that has to drop to a smaller model.
| Memory | What it comfortably runs at 4-bit |
|---|---|
| 16GB | Small models, 3B to 8B |
| 24GB | Mid-size models, up to roughly the 14B class |
| 32GB | A 20B model, with room for context |
These pairings follow from the 0.6 rule, so treat them as starting points rather than guarantees. Your own context length and the apps you keep open will move the line.
Key takeaways
- Unified memory is the ceiling, not the GPU, and a 32GB Mac will run a 20 billion parameter model at 4-bit while a 16GB Mac will not.
- Multiply the parameters in billions by 0.6 to estimate the gigabytes needed at 4-bit, so a 14B model wants about 8.4GB before the context window is added.
- Q4_K_M quantisation is the usual sweet spot, since 8-bit roughly doubles the memory for a quality gain most people cannot reliably detect.
- Ollama is the fastest way in, LM Studio offers a GUI and local API server, and llama.cpp exposes every setting.
- Anything needing current information or a frontier-scale model still belongs in the cloud, while local wins on privacy, latency and having no monthly bill.