Future TechnologyFuture Technology
AI

Self-Host an LLM or Rent an API? The Decision, Without the Hype

· 2 min read · By Future Technology

Key takeaways

  • Platform data from mixed-workload providers shows blended token costs falling 40 to 65 percent when suitable work moves to open weight models
  • Rent for frontier reasoning, spiky volume and no infrastructure staff; own for high predictable volume, latency limits and data that cannot leave your estate
  • The claim that half the Fortune 500 uses open source models is a vendor figure we could not verify independently

Platform data from providers routing mixed workloads shows blended token costs falling 40 to 65 percent when suitable work moves to open weight models. That is the number worth holding onto when you decide whether to self host an LLM. Most of the rest of this argument is vendor positioning.

When renting is the right answer

Rent when you need frontier reasoning that smaller models cannot do. Low or spiky volume points the same way, because an idle GPU costs the same as a busy one and utilisation is the whole economic case for owning hardware. The other two conditions are staffing and sensitivity: no infrastructure people, and data that does not justify the operational overhead of keeping it in house.

The frontier gap is real in narrow places and overstated in most. Voice and multimodal work is one of the places it still holds, which is part of why live voice models stayed hosted long after text got commoditised.

When owning wins

Own when volume is high and predictable, the condition that makes fixed hardware cheaper per token than metered access. Latency is the second case, where a round trip to someone else's region is the bottleneck, and legal constraints on where data can sit are the third. The last one is the most underrated: if a task is narrow enough that a small fine-tuned model matches a frontier one, owning it is cheap, and that covers more tasks than most teams assume before they measure.

Hardware is the caveat. Memory is the binding constraint on inference serving, and current pricing reflects that, as our explainer on HBM4 memory sets out. A build costed at last year's prices does not survive contact with this year's quotes.

The middle path most teams land on

Route by task. Cheap open weight models handle classification, extraction, summarisation and routing. Frontier APIs handle the genuinely hard reasoning, which is also where vendor-integrated products like Salesforce and Anthropic's tie-up are aimed. That split is where the 40 to 65 percent comes from, not from replacing everything with something you host yourself.

Hugging Face's chief executive has claimed roughly half the Fortune 500 now uses open source models alongside rented APIs. Treat the figure carefully; it is a vendor claim and we could not verify it independently. The direction is well evidenced even where the percentage is not.

So: measure which of your calls actually need frontier reasoning before you price any of this. In most estates the share is smaller than the bill implies, and that measurement is the entire decision.

Get the briefing, free

The biggest tech story, explained in 3 minutes every weekday. Choose your briefings →

Free. No spam. Unsubscribe in one click.