Future TechnologyFuture Technology
AI

OpenAI's Jalapeno Chip Beat Nvidia's Best, and OpenAI Is Still Buying Nvidia

· 4 min read · By Future Technology

Key takeaways

  • Jalapeno delivered 1.5x to 1.9x more AI work per watt and 1.7x to 3.6x lower end to end latency than the best commercially available systems
  • On interactive workloads, where a human is waiting on tokens, the latency gap widens to 2.1x to 4.1x
  • Benchmarks ran on SemiAnalysis's public InferenceX suite with some runs verified on site, a higher bar than a vendor slide
  • Only a very small Jalapeno deployment lands by the end of 2026, with the real rollout in 2027, and OpenAI has committed to keep buying Nvidia

The number to hold onto is 1.9. That is how many times more AI work per watt OpenAI says its own inference chip does at peak throughput, measured against the best systems you can currently buy.

Jalapeno is OpenAI's custom inference silicon, and this week the company published the first real benchmarks for it. Across three models, GPT-OSS 120B, DeepSeek R1 670B and Kimi K2.5 1T, it claims 1.5x to 1.9x more work per watt and 1.7x to 3.6x lower end to end latency. On interactive workloads, the kind where you are sitting there watching tokens appear, the latency advantage stretches to 2.1x to 4.1x.

What makes the OpenAI Jalapeno chip numbers unusual

Most custom silicon announcements are a slide deck. This one ran on SemiAnalysis's public InferenceX benchmark, and SemiAnalysis verified some of the runs on site. That does not make the results beyond question, but it is a meaningfully higher bar than a vendor publishing its own graph and asking you to trust the axis labels.

The caveat matters as much as the headline, and SemiAnalysis raised it themselves. The fair comparison is not Nvidia's Blackwell but Vera Rubin, because both Jalapeno and Vera Rubin use HBM4 memory. Against Vera Rubin the gap narrows considerably, though Jalapeno still edges ahead on output tokens per megawatt. It does that without multi token prediction, an optimisation Nvidia already ships and OpenAI has not adopted yet. Whether that headroom is real or theoretical is the open question.

Why inference efficiency is now the whole game

Training a frontier model is a capital event that happens a handful of times. Inference happens every time anyone types anything, forever. That flips the economics: a watt saved per token is money that stops flowing to Nvidia, multiplied by billions of requests a day.

This is the same shift that has been reshaping the accelerator market all year. Nvidia cancelled Rubin CPX as Groq's LPX landed, and the reason was the same one: inference workloads have different shapes to training workloads, and purpose-built silicon can exploit that. If you want the underlying distinction between the processor types doing this work, our explainer on NPU vs GPU vs CPU covers what each one is actually good at.

The gap between the benchmark and the truck

Here is the part that gets skipped in the excitement. OpenAI has committed to keep buying Nvidia hardware. Only a very small Jalapeno deployment arrives by the end of 2026, with the real rollout scheduled for 2027.

That is not a contradiction, it is the actual story. Designing a chip that wins a benchmark and manufacturing millions of them that work reliably in a data centre are entirely different problems, separated by fab capacity, packaging supply, HBM4 allocation and years of software maturity. Nvidia's moat was never only the silicon. It was CUDA, the tooling, and the fact that a rack of it turns up when you order one.

So the honest read is this: OpenAI has proved it can design competitive inference silicon, which changes its negotiating position enormously. It has not yet proved it can replace Nvidia, and its own purchase orders say it knows that. The company's Astra model quietly clearing unsolved maths problems shows where the demand for all those watts is coming from.

Watch the 2027 deployment figures, not the benchmark chart. That is where this either becomes real or stays a very expensive proof of concept.

Get the briefing, free

The biggest tech story, explained in 3 minutes every weekday. Choose your briefings →

Free. No spam. Unsubscribe in one click.