Hardware

vLLM vs Ollama is the wrong question if you own spare machines

(4 days ago) · 4 min read · By Future Technology

Key takeaways

  • vLLM is built for throughput and concurrent requests, Ollama for getting one model running in a single command.
  • LocalAI exists to swap a cloud provider out without rewriting client code, because it speaks the same API.
  • exo pools laptops, desktops and spare machines into one inference target, which changes the hardware maths entirely.
  • Pick the layer by how many people will use it and what hardware you already own, not by benchmark charts.

A Hacker News thread about running frontier models on hardware you own pulled 138 points and 71 comments this week. Very little of it was about GPUs. People had already worked out the hardware and then stalled on the software that actually serves the model.

The question tends to get framed as vLLM vs Ollama, which is the wrong framing. There are four serious options in the local inference stack, and each one answers a different problem.

vLLM vs Ollama: throughput against getting started

vLLM is the throughput answer. It was built to serve many concurrent requests efficiently, with paged attention and continuous batching doing the work that naive serving wastes. If a model sits behind an internal tool that thirty people hit at once, this is the one. The setup cost is real and so is the payoff.

Ollama is the other end of the range. One command pulls a model and starts serving it, and the defaults are sensible enough that most people never touch them. For a single user on a single machine, running vLLM instead is work you did not need to do.

The split is concurrency rather than quality. One person asking questions gets the same answer from either, and thirty people at once do not.

LocalAI is the compatibility layer

LocalAI solves a narrower problem. You have code written against a cloud provider, and you want to stop paying per token without rewriting the client. It speaks the same API, so the change is a base URL and a key.

That is a smaller promise than the other two make, and it is the one that saves the most engineering time when it fits. If your stack already talks to a hosted endpoint, this is the shortest path off it.

exo uses the machines you already have

exo is the odd one out. Rather than asking what a single box can hold, it spreads a model across whatever devices are on your network. A laptop, a desktop and an old machine in a cupboard become one inference pool.

That reframes the hardware question entirely. The usual VRAM maths for a local model assumes one machine has to fit the whole thing. exo assumes it does not, which puts a 70B model within reach of people who were never going to buy a second 24GB card.

The tradeoff is latency and fragility. Network hops between devices cost you tokens per second, and every extra machine is one more thing that can drop off mid-generation.

How to pick

Start with users rather than benchmarks. One person on one machine wants Ollama, and a team hitting a shared endpoint wants vLLM. If you already have code aimed at a cloud API, LocalAI is the shortest move off it. If you have several underused machines and no budget for a bigger one, exo is the only option here that helps.

Then check what the hardware does under load. A dual-GPU build draws more than most people expect, and a serving box that falls over during a power dip is worse than no serving box, which is the case for a UPS on a home server. System RAM matters more than people assume once layers start offloading, and a 64GB DDR5 kit costs less than the GPU you were trying to avoid buying.

The part worth watching is 15 October, when StepFun opens the weights on a 600B model. Whatever serving layer you settled on is about to be tested by a much larger file.

Some links in this article are affiliate links. We may earn a small commission at no extra cost to you.

More from Future Technology