The Numbers Nvidia Has To Answer
OpenAI unveiled its first custom AI chip, code-named Jalapeno, this week, and the accompanying benchmark numbers are the real story. Using the public InferenceX benchmark, independently verified in part by the analysis firm SemiAnalysis, Jalapeno produced 1.5 to 1.9 times more AI work per watt at peak throughput than Nvidia's currently shipping Blackwell Ultra (GB300) systems, and delivered 1.7 to 3.6 times lower end-to-end latency in tests running the DeepSeek R1, Kimi K2.5 and GPT-OSS models.
| Chip | Peak compute | Package power | Status |
|---|---|---|---|
| OpenAI Jalapeno | 13.4 PFLOPs (MXFP4) | 700W | Engineering samples only |
| Nvidia Blackwell Ultra (GB300) | Not disclosed at same precision | 1,400W | Shipping now |
| Nvidia Vera Rubin | 17.5 PFLOPs | 900-1,150W | Shipping to customers |
OpenAI says the chip took roughly 16 months from design start in mid-2024 to final fabrication in November 2025, with earlier OpenAI models assisting the chip design itself and newer ones accelerating the programming work.
The Comparison Even Its Author Calls Unfair
What makes this launch different from a routine benchmark war is that OpenAI and the analysts who verified the results both say the headline Blackwell comparison undersells the real picture. Jalapeno runs on HBM4 memory with 15.4 terabytes per second of bandwidth, a newer memory generation than Nvidia's currently shipping Blackwell line uses. SemiAnalysis noted the fairer rival is Nvidia's Vera Rubin, which also uses HBM4 and already ships 17.5 PFLOPs at 900 to 1,150 watts to customers today, while Jalapeno remains an engineering sample with no shipping date.
Against Rubin, the two chips land at a comparable cost per token, but the methodology is not identical. Jalapeno's numbers reflect single-token prediction, while Rubin's tests used speculative decoding, a technique that changes throughput without changing which chip is faster in principle. The benchmark suite that would test longer, production-representative context windows has not been run yet.
OpenAI's Own CFO Downplayed The Replacement Story
The most telling reaction to the launch did not come from Nvidia or from outside analysts. It came from OpenAI's own chief financial officer, Sarah Friar, who described Jalapeno as something that "fits into a broader compute strategy where data centers, chips, models, the developer platform, products, and devices all work as one integrated system," and said explicitly that the chip "complements OpenAI's existing partnerships" rather than replacing them.
OpenAI has not withdrawn from any of its compute relationships. The company continues to build capacity with Nvidia, AMD, AWS, Cerebras and CoreWeave, and has made no statement suggesting an intent to reduce its dependence on any of them. Jalapeno reads less like a competitor's product launch and more like an insurance policy against future supply constraints, at a moment when every major compute buyer still says demand outstrips available hardware.
What This Means For Anyone Buying AI Compute
The practical lesson for a business planning its own AI infrastructure spend is not that Nvidia has a new problem. It is that a benchmark headline from a compute buyer who is now also a compute maker is not the same signal as a genuine market alternative arriving at scale. Jalapeno exists in engineering-sample quantities. It has no announced price, no announced availability date for outside customers, and OpenAI itself frames it as a hedge rather than a replacement.
Any procurement decision that treats this week's numbers as a reason to renegotiate a Nvidia contract, or to bet on a near-term multi-vendor inference market, is reading further into the announcement than OpenAI itself is willing to claim. The chip is real, the efficiency gain is real, and the caveats attached to the comparison by OpenAI's own launch partner are just as real.
Read next: Samsung's Quarter Set to Beat Nvidia on Profit | Ohio Ratepayers Won't Pay for OpenAI's 8GW Campus



