OpenAI's Jalapeno chip beat Nvidia's best, and OpenAI is still buying Nvidia
Key takeaways
- Jalapeno delivered 1.5x to 1.9x more AI work per watt and 1.7x to 3.6x lower end to end latency than the best commercially available systems
- On interactive workloads, where a human is waiting on tokens, the latency gap widens to 2.1x to 4.1x
- Benchmarks ran on SemiAnalysis's public InferenceX suite with some runs verified on site, a higher bar than a vendor slide
- Only a very small Jalapeno deployment lands by the end of 2026, with the real rollout in 2027, and OpenAI has committed to keep buying Nvidia
The number to hold onto is 1.9. That is how many times more AI work per watt OpenAI says its own inference chip does at peak throughput, measured against the best systems you can currently buy.
Jalapeno is OpenAI's custom inference silicon, and this week the company published the first real benchmarks for it. Across three models, GPT-OSS 120B, DeepSeek R1 670B and Kimi K2.5 1T, it claims 1.5x to 1.9x more work per watt and 1.7x to 3.6x lower end to end latency. On interactive workloads, the kind where you are sitting there watching tokens appear, the latency advantage stretches to 2.1x to 4.1x.
Inference means running a trained model to produce answers, as opposed to training it. A token is the small chunk of text a model reads or writes at a time, and latency is how long you wait for them. Work per watt matters because electricity and cooling are now among the biggest costs of running a data centre.
What makes the OpenAI Jalapeno chip numbers unusual?
Most custom silicon announcements are a slide deck. This one ran on SemiAnalysis's public InferenceX benchmark, and SemiAnalysis verified some of the runs on site. That does not make the results beyond question, but it is a meaningfully higher bar than a vendor publishing its own graph and asking you to trust the axis labels.
The benchmark covers three very different models. GPT-OSS 120B is a mid-sized open-weight model, DeepSeek R1 670B is a much larger reasoning model, and Kimi K2.5 1T sits at roughly a trillion parameters. A chip that holds its lead across that spread is harder to dismiss as a result tuned for one workload.
Benchmarks like InferenceX matter because inference results are easy to flatter. Batch size, model precision and the mix of short and long prompts can all swing a result, so a public suite with fixed rules and outside checking narrows the room for selective reporting. The figures are still OpenAI's claim, and independent repeat runs on shipping hardware would carry more weight.
Is Jalapeno really faster than Nvidia's best?
The caveat matters as much as the headline, and SemiAnalysis raised it themselves. The fair comparison is not Nvidia's Blackwell but Vera Rubin, because both Jalapeno and Vera Rubin use HBM4 memory. HBM4 is the stacked, high bandwidth memory that sits next to the processor and feeds it data, and for inference it often matters more than raw compute.
Against Vera Rubin the gap narrows considerably, though Jalapeno still edges ahead on output tokens per megawatt. It does that without multi token prediction, an optimisation Nvidia already ships and OpenAI has not adopted yet. Multi token prediction lets a system guess several tokens ahead and check them in one pass, which raises throughput without new hardware. Whether that headroom is real or theoretical is the open question.
| Claim | What was reported | Caveat |
|---|---|---|
| Work per watt | 1.5x to 1.9x versus the best commercial systems | Narrower against Vera Rubin |
| End to end latency | 1.7x to 3.6x lower | Measured on three named models |
| Interactive latency | 2.1x to 4.1x lower | Applies where a human waits on tokens |
| Verification | SemiAnalysis InferenceX, some runs checked on site | Not every run was verified on site |
| Deployment | Very small by end of 2026, real rollout in 2027 | OpenAI still buying Nvidia |
Why does inference efficiency matter so much now?
Training a frontier model is a capital event that happens a handful of times. Inference happens every time anyone types anything, forever. That flips the economics: a watt saved per token is money that stops flowing to Nvidia, multiplied by billions of requests a day.
This is the same shift that has been reshaping the accelerator market all year. Nvidia cancelled Rubin CPX as Groq's LPX landed, and the reason was the same one: inference workloads have different shapes to training workloads, and purpose-built silicon can exploit that. Our piece on why Nvidia won training but inference is a different problem lays out how that split opened the door to rivals.
If you want the underlying distinction between the processor types doing this work, our explainer on NPU vs GPU vs CPU covers what each one is actually good at. For the narrower question of how the chips themselves differ, see inference chips versus training chips.
Why is OpenAI still buying Nvidia?
Here is the part that gets skipped in the excitement. OpenAI has committed to keep buying Nvidia hardware. Only a very small Jalapeno deployment arrives by the end of 2026, with the real rollout scheduled for 2027.
That is not a contradiction, it is the actual story. Designing a chip that wins a benchmark and manufacturing millions of them that work reliably in a data centre are entirely different problems, separated by fab capacity, packaging supply, HBM4 allocation and years of software maturity. Nvidia's moat was never only the silicon. It was CUDA, the tooling, and the fact that a rack of it turns up when you order one.
Software is the quiet part of that moat. Every model, library and serving stack OpenAI runs today has been tuned for Nvidia's tools over years. Moving a production workload to a new architecture means rewriting kernels, retesting reliability and training engineers, and none of that shows up on a benchmark chart.
What happens next?
So the honest read is this: OpenAI has proved it can design competitive inference silicon, which changes its negotiating position enormously. It has not yet proved it can replace Nvidia, and its own purchase orders say it knows that.
Even a partial shift matters. If a customer of OpenAI's size can credibly run a share of its inference on its own chips, Nvidia has to price against that option. The pressure shows up in contract terms and delivery slots well before it shows up in market share.
For readers, the practical effect is indirect. Cheaper inference tends to show up as lower prices for model access, faster responses in chat tools and fewer usage limits, though none of that is guaranteed to be passed on. The people who feel it first are developers paying per token.
The other thing to track is whether OpenAI adopts multi token prediction on Jalapeno, since that is the clearest sign the remaining headroom is real.
Watch the 2027 deployment figures, not the benchmark chart. That is where this either becomes real or stays a very expensive proof of concept.
Key takeaways
- OpenAI's Jalapeno chip delivered 1.5x to 1.9x more AI work per watt and 1.7x to 3.6x lower end to end latency than the best commercially available systems.
- On interactive workloads, where a human is waiting on tokens, the latency advantage widens to 2.1x to 4.1x.
- The benchmarks ran on SemiAnalysis's public InferenceX suite, with some runs verified on site, which is a higher bar than a vendor slide.
- Against Nvidia's Vera Rubin, which also uses HBM4, the gap narrows considerably, though Jalapeno still edges ahead on output tokens per megawatt.
- Only a very small Jalapeno deployment lands by the end of 2026, with the real rollout in 2027, and OpenAI has committed to keep buying Nvidia.