STANFORD, 27 AUG 2026 — OpenAI has published the first benchmarks for Jalapeño, its own inference accelerator, and they are better than Nvidia's Blackwell on the measure the industry now cares most about. Peak performance per kilowatt came in at 85,448 against 44,960 mixed tokens per second, roughly 1.9 times.

The chip was presented at Hot Chips, the same conference where Nvidia detailed its Vera processor. OpenAI says it will deploy Jalapeño in its own infrastructure by the end of the year, and that it will go on buying Nvidia hardware.

What was shown

Across the tested range, Jalapeño delivered 1.5 to 1.9 times more AI work per watt than Blackwell and 1.7 to 3.6 times lower latency. On highly interactive workloads, the short conversational requests a chat product generates constantly, it ran 2.1 to 4.1 times faster.

That last category is the one OpenAI built for, and the spread of the results tells you as much as the peak figure does. A chip that is uniformly 1.9 times better would suggest a general advantage. One that ranges from 1.5 to 4.1 depending on the workload is a chip tuned very precisely for a particular shape of traffic.

1.9xPeak performance per kilowatt
1.7-3.6xLower latency
2.1-4.1xOn interactive workloads
End of 2026Deployment in OpenAI's own fleet

A specialist beat a generalist at the specialist's job

The comparison needs its terms stated, because the headline invites a conclusion the numbers do not support.

Blackwell is a general-purpose accelerator. It trains models and it serves them, it runs workloads nobody at Nvidia has seen, and it has to be good at all of that for customers whose requirements differ wildly. Jalapeño does one thing: serve OpenAI's models to OpenAI's traffic.

Winning on inference efficiency against a part that also has to train is the expected result of that trade. It is what specialisation buys, and it is why Google, Amazon and now OpenAI have all built their own. The more interesting question, how Jalapeño compares against Nvidia's own inference-oriented silicon on the same benchmark, is not one these figures answer.

This does not make the achievement small. Designing a competitive accelerator is extremely hard, and the margin is large enough to change OpenAI's cost structure. But it is not the same claim as beating Nvidia.

The benchmark deserves more credit than the number

The measurement was taken with SemiAnalysis's InferenceX, a public third-party benchmark that measures the full request pipeline and reports throughput, power draw and latency together rather than separately.

That is a higher standard than the industry norm. Vendor-defined benchmarks reported in isolation are how a chip claims a throughput record while drawing twice the power, or a latency record at a batch size nobody runs in production. Measuring all three at once removes most of the room for that.

The remaining caveat is the ordinary one. OpenAI ran the benchmark, on its own hardware, and no independent party has reproduced it. Publishing against a public methodology means an independent party could reproduce the result. That is more than most such announcements offer, including the Vera figures Nvidia presented at the same conference against its own internal comparisons.

Per-watt is the metric because power is the constraint

A few years ago the headline number would have been throughput. That it is now performance per watt reflects what actually limits an AI business.

Capacity is bounded by electricity. A campus has a grid connection measured in megawatts, that connection took years to secure, and no amount of capital shortens the queue. Within that fixed envelope, a chip delivering 1.9 times the work per watt does not merely reduce the power bill — it increases how much serving the site can do at all.

That converts an engineering result into a strategic one. If OpenAI can serve nearly twice the traffic from the same substation, it has effectively doubled its capacity without a permit, a transformer or a single new megawatt. For a company whose expansion is gated by power, that is worth more than the silicon cost saving.

OpenAI is Nvidia's customer, and Nvidia is OpenAI's guarantor

The relationship between the two companies is unusual.

OpenAI has published benchmarks showing its chip outperforming Nvidia's, days before Nvidia reports quarterly results. It has also said it will keep buying Nvidia hardware broadly, for training and inference both. And Nvidia has agreed to guarantee up to US$105bn of OpenAI's Ohio data centre lease, with itself as exclusive chip provider for that site.

Those facts are not contradictory, and the combination is the point. A large buyer that can credibly build its own parts negotiates differently from one that cannot, and the announcement is worth something to OpenAI whether or not Jalapeño ever displaces much Nvidia hardware. Publishing the numbers is itself a commercial act.

It is also a reminder of how concentrated this market's relationships are. The supplier is underwriting the customer's property lease while the customer publishes benchmarks against the supplier's product, and both companies are behaving rationally.

The hard part comes after the benchmark

A working chip with good numbers is perhaps a third of the problem. What remains has defeated better-funded efforts than this one.

It has to be manufactured in volume, competing for leading-edge capacity and advanced packaging against everyone else. It needs a software stack that the people deploying models will actually use — Nvidia's moat, built over two decades, which has drowned several technically sound competitors. And it must bet on the model architectures of 2026, in a field that has not stayed still for five years.

OpenAI has one advantage that mitigates all three. It only has to satisfy itself. A merchant vendor must support every customer's model; OpenAI needs Jalapeño to run OpenAI's models well, and it controls both sides.

What it means from here

For operators in the region the useful figure is not 1.9. It is that inference efficiency has become the axis of competition, which changes what a compute contract should be measured on.

Capacity bought in tokens per second per megawatt is a different purchase from capacity bought in accelerators, and the gap between the two is now large enough to matter to a budget. Anyone signing for inference capacity should be asking what the provider's serving efficiency actually is, because it is no longer a detail of their operations — it is most of what determines the price they can offer.

The second-order effect is on Nvidia's pricing power rather than its volumes. Every large buyer that demonstrates a credible alternative for its own dominant workload changes the negotiation, and the demonstration matters even when the buyer keeps buying.