AI

OpenAI cut GPT-6 pricing in half and called it permanent

(2 days ago) · 4 min read · By Future Technology

Key takeaways

  • GPT-6 Sol runs at 2 dollars per million input tokens and 10 dollars output, against 4 and 20 for GPT-5.6 Sol.
  • Luna drops to 10 cents input and 50 cents output, with a further 90 percent discount on cached input reads.
  • OpenAI says the cut is permanent list pricing driven by caching and inference gains, not a subsidised promotion.
  • Agent loops that were marginal at 20 dollars per million output tokens now clear their economics comfortably.

Ten cents per million input tokens. That is what GPT-6 Luna costs as of 22 September, down from 20 cents, and it is the cheaper half of a GPT-6 pricing change that takes 50 to 60 percent off OpenAI's rates across the new lineup.

OpenAI released Sol and Luna on 22 September, sitting either side of the flagship Astra. Sol runs at 2 dollars per million input tokens and 10 dollars per million output, against 4 and 20 for GPT-5.6 Sol. Luna lands at 10 cents in and 50 cents out, down from 20 cents and 1.20 dollars. Cached input reads take a further 90 percent discount on top of those numbers.

Where GPT-6 pricing sits now

The two models are aimed at different jobs. Sol is pitched at repeated technical work, the kind that involves building a feature, reviewing the code and then debugging what broke. Luna is the high volume everyday tier, the model you point at a million support tickets rather than one hard problem.

The cached input discount is the part that gets missed. If your prompt carries a large fixed context, a system prompt, a schema, a chunk of documentation, you pay full rate once and a tenth of it afterwards. Applications with heavy fixed context see a bigger effective cut than the headline number suggests.

Why permanent is the word that matters

An OpenAI spokesperson confirmed these are permanent list prices rather than promotional rates, and attributed the drop to caching and inference improvements rather than subsidy. That distinction decides whether anyone can plan around it.

Promotional pricing is a marketing line. You cannot build a product on it, because the rate can revert at a quarter's notice and your margin goes with it. Permanent list pricing is an input cost, and input costs go into spreadsheets. The companies raising serious money on coding agents, including Cognition at a 2 billion dollar Series E, have been modelling against the old numbers.

What it does to agent economics

An agent loop that reads a codebase, proposes a change, runs the tests and retries burns output tokens on every pass. Halving the output cost does not make the loop twice as good. It makes roughly twice as many passes worth running before the bill stops making sense, which is a different and more useful thing.

There is a knock-on effect for local inference. Running a model on your own hardware has always been part cost argument, part control argument. At 10 cents per million tokens the cost half gets thinner, which leaves privacy, latency and not depending on someone else's uptime doing the work. The case for vLLM, Ollama and the rest of the local stack is still there, it just rests on fewer legs. Open weight releases like StepFun's Step 5 preview compete on the same ground with less price cover than they had a week ago.

What to watch

Anthropic and Google now have to answer a 50 percent cut from the largest API vendor, and the obvious answer is a matching cut of their own. The number worth watching is not the next price announcement. It is whether OpenAI's claim about inference efficiency holds up, because a cut funded by real cost reduction survives a bad quarter and a cut funded by market share does not.

More from Future Technology