AI

Step 5 Preview benchmarks put it seven points behind GPT-6 Astra Max

(4 days ago) · 3 min read · By Future Technology

Key takeaways

  • Step 5 Preview is a 600B sparse mixture of experts model with 27B active parameters and a one million token context window.
  • It scores 67.7 on DeepSWE v1.1 against 74.1 for GPT-6 Astra Max, so the gap to the closed frontier is roughly seven points.
  • Open weights are scheduled for 15 October, and pricing sits near one dollar per million input tokens.

StepFun announced Step 5 Preview on 20 September. It is a sparse mixture of experts model with 600 billion total parameters, 27 billion of them active at any one time, and a one million token context window. Open weights are scheduled for 15 October.

Step 5 Preview benchmarks show a very tight race at the top of the open field. On DeepSWE v1.1 at high reasoning it scores 67.7, against 67.5 for Kimi K3 Max and 66.9 for GLM-5.3 Max. The closed models still sit above it, with GPT-6 Astra Max at 74.1 and Claude Opus 5 Max at 74.0. On Terminal-Bench 4.0 it reaches 33 percent, ahead of DeepSeek V4.1 Flash at 27 percent.

The quiet number is the price

Artificial Analysis puts Step 5 Preview at 44 on its Intelligence Index, where the median for reasoning models in the same price tier is 24.

Pricing lands near one dollar per million input tokens. For long-horizon agent work, where one run can chew through an enormous context several times over, cost per run is what decides whether a workflow ships or stays a demo. Context length and price are doing more work here than the benchmark table is.

Why 15 October matters more than the scores

A seven-point gap to the closed frontier is real, and it is also smaller than the gap a year ago. Open weights change the calculation in a way a cheaper API does not, because you can run the model on hardware you control. That removes the per-token bill and the question of where the data goes, and replaces both with the problem of picking a serving layer that can handle a 600B sparse model.

The number worth watching is whether the released weights reproduce the preview scores. Benchmarks published before a public release are run by the people who built the model, and the independent reproductions in the week after 15 October will say more than anything announced now.

The second thing to watch is pricing at the top. The capital going into frontier training, from OpenAI's projected cash burn to the share of engineering that Claude now contributes to building the next Claude, assumes a durable gap. Seven points, closing, with the weights free next month, is a harder assumption to fund.

More from Future Technology