GPT-6 Astra Benchmarks: OpenAI Ships the Model It Paused
Key takeaways
- GPT-6 Astra launched on 3 September across every ChatGPT tier, with a 1M token context window and the API ID gpt-6-astra
- It scores 72.6 percent on OSWorld 2.0 offline, ahead of Claude Opus 5 at 70.2 and GPT-5.6 Sol at 65.7
- API pricing is 10 dollars per million input tokens and 50 per million output, roughly 2.5 times Sol's promotional rate
- OpenAI paused the model four weeks ago over cyber capability, and it now scores 100 percent on ExploitBench
Four weeks ago OpenAI publicly slowed GPT-6 Astra down because it was getting too good at cyber. On 3 September the model shipped, and it scores 100 percent on ExploitBench. Those two facts belong in the same paragraph.
Astra is now live across every ChatGPT tier including Plus, and available on the API under the model ID gpt-6-astra. It was pre-trained on more than 100,000 GPUs at Stargate, takes multimodal input, and carries a 1M token context window.
The GPT-6 Astra benchmarks that matter
The headline numbers are high even by launch-day standards. Astra posts 97.6 percent on FrontierMath Tier 4 and 99.9 percent on ARC-AGI-3, the latter measured under OpenAI's own provider adapter harness. That caveat is worth printing rather than burying.
The more useful figure is OSWorld 2.0, which tests agentic computer use rather than reasoning in isolation. On the offline set Astra takes 72.6 percent, against 70.2 for Claude Opus 5 and 65.7 for GPT-5.6 Sol. That is a 2.4 point lead over the nearest model.
What it costs
API pricing is 10 dollars per million input tokens and 50 per million output. Cached input drops to 1 dollar, batch runs at half price, and Fast mode doubles the rate. Against Sol's current promotional pricing that is a 2.5 times increase, and it puts Astra roughly level with Anthropic's Fable 5.1.
So the arithmetic for anyone buying tokens is 2.4 points of OSWorld for 2.5 times the price. On agent workloads running all day that gap has to earn its keep. On a one-off summary it does not, which is why the cached input tier at 1 dollar is the line most teams should read first.
The pause is the story
In early August OpenAI said it was holding Astra back over cyber capability concerns, and we covered the decision to pause at the cyber threshold at the time. The model that came out of that pause saturates the security benchmark it was paused over.
That can mean one of two things. Either the mitigation work is good enough that a perfect ExploitBench score no longer implies uncontrolled capability, or ExploitBench stopped being a useful ceiling somewhere between August and September. OpenAI has not published enough for anyone outside the company to tell which.
It is the same shape as Astra clearing open mathematics problems earlier in its testing cycle. The capability lands publicly with a number attached, and the safety reasoning arrives later as a summary.
What to watch
Frontier pricing has moved up rather than down for the first time in a while, and the company moving it is the one filing to go public. Watch whether Anthropic and Google hold their current rates through the autumn. If they hold, Astra's price is a bet that the benchmark gap is worth paying for. If they follow, the run of frontier models getting cheaper every quarter is finished.