The numbers OpenAI published

OpenAI used the Hot Chips 2026 conference on 25 August to release the first public benchmark results for Jalapeno, the inference chip it built with Broadcom in about nine months. Run on SemiAnalysis's InferenceX benchmark across three models, GPT-OSS-120B, DeepSeek R1 and Kimi K2.5, Jalapeno delivered 1.5 to 1.9 times more work per watt than Nvidia's GB300 and cut end-to-end response latency by 1.7 to 3.6 times, rising to 2.1 to 4.1 times on the most interactive workloads.

The chip carries a 700-watt power rating but drew at or below 550 watts across the tested workloads, and a single 128-chip rack delivers 1.7 exaflops of 4-bit compute alongside 27.5 terabytes of HBM4 memory, with a full pod scaling to 2,048 chips. OpenAI plans small-volume deployment by the end of 2026, with scaling through 2027.

MetricOpenAI's claim
Work per watt1.5x to 1.9x more than Nvidia GB300
End to end latency1.7x to 3.6x lower than Nvidia GB300
Rated power draw700 watts, measured at or below 550 watts in tests
Deployment timelineSmall volume by end of 2026, scaling through 2027

Why a power-per-watt number matters more in Europe than the headline suggests

A genuine efficiency gain of this size would matter well beyond OpenAI's own data centres, because it lands in the middle of Europe's live argument about whether the grid can support the AI buildout at all. Ireland's grid operator has frozen new data-centre connections around Dublin until 2028, and this week's reporting on Hungary's spot-power spikes shows how tight European interconnection capacity already is; a chip that genuinely needs 40 percent less power per unit of AI work would ease exactly that bottleneck.

The gap is that nobody outside OpenAI has run the comparison yet. The benchmark used OpenAI's own model selection, OpenAI's own choice of a GB300 comparison rather than Nvidia's newest Vera Rubin generation, and no third-party lab result has been published alongside it.

What to do with a vendor's own number

Treat this the way any procurement team should treat a vendor-run benchmark: real information, worth tracking, not yet a planning figure. Independent efficiency data typically arrives through MLPerf inference results or matched third-party lab tests run on hardware the vendor does not control, and neither exists yet for Jalapeno.

For any EU infrastructure planner, utility, or grid operator fielding a connection application that cites next-generation chip efficiency as justification for capacity, the correct response is to ask which benchmark backs the number and who ran it, not to accept a vendor's own slide as a substitute for an independent test.