OpenAI Jalapeño vs Nvidia: What the Benchmarks Mean
OpenAI posted first Jalapeño benchmarks: up to 1.9x more throughput per kilowatt than Nvidia. Real numbers, plus the caveats the headlines are skipping.
On August 25, 2026, OpenAI put numbers next to a name it had only teased before. Jalapeño, its first custom inference chip, built with Broadcom, ran against Nvidia’s GB200 and GB300 rack systems on SemiAnalysis’s public InferenceX benchmark, with OpenAI’s own engineers in the room. The result: 1.5x to 1.9x more throughput per kilowatt, and 1.7x to 3.6x lower end-to-end latency, across three open-weight models.
Most of the coverage since has been a specs recap: die size, HBM4 bandwidth, watts. That is the wrong unit of analysis. The number that actually matters to anyone thinking about AI infrastructure spend is watts per rack, because at gigawatt scale, power is the thing OpenAI cannot buy its way out of. Jalapeño is not really a chip story. It is a capacity story, and a bet that vertical integration beats buying merchant GPUs from Nvidia, at least for the one workload OpenAI runs more of than anyone alive: serving its own models to its own users.
Why perf-per-watt is the currency that matters
Ask any hyperscaler what actually caps how much AI compute it can deploy this year and the answer is rarely money. It is power. A gigawatt-scale AI data center already runs Nvidia’s GB300 NVL72 racks at roughly 135 kW each, an order of magnitude above a conventional data center rack, and the grid interconnection queue to add another gigawatt of capacity can run years, not quarters. When the binding constraint is watts, not dollars, the chip that does more work per watt effectively buys more compute than the chip that is merely faster.
That is the frame Jalapeño’s numbers were built for. Each compute die draws 700W against Nvidia parts that run from roughly 900W up to 1,800W in max-throughput configurations, and OpenAI’s own rack packs 128 Jalapeño ASICs into about 160 kW total, split across sixteen host trays and sixteen accelerator trays. Fewer watts per unit of useful output means more racks fit inside a given power budget, which means more served requests inside a given interconnection agreement. That is the actual economics story: not “faster chip,” but “more inference inside the same substation.”
| Item | High end of reported range | Low end of reported range |
|---|---|---|
| Throughput per kilowatt | 1.9x | 1.5x |
| End-to-end latency reduction | 3.6x | 1.7x |
At the narrowest, most favorable slice of the comparison, OpenAI reports 8.6x to 104.3x more throughput per kilowatt. That number is real, but read the fine print: it holds only at Jalapeño’s fastest, lowest-latency operating point measured against GB300’s fastest previous time-between-tokens setting, not the general comparison above. It is a genuinely large gap in one narrow configuration, not the headline number, and treating it as the latter overstates the case.
The vertical-integration bet: Broadcom over merchant silicon
Nvidia’s business model sells the same GPU to everyone: OpenAI, Microsoft, a mid-sized neocloud, a crypto miner repurposing rigs. That horizontal model is enormously profitable and it is also, by construction, a compromise. A chip built to run every workload well cannot be built to run one workload optimally. Jalapeño is OpenAI’s bet that the second path, in Broadcom’s words a full custom design, wins for the one job OpenAI does at planet scale: inference on its own models.
This is not a new playbook. Google built its first Tensor Processing Unit in 2015 as an inference-only chip for Search, and TPUs now also serve Gemini and much of Google Cloud’s internal inference workload a decade later. Amazon’s Trainium and Inferentia chips exist for exactly the same reason: a hyperscaler’s own workload is predictable and enormous enough to justify a chip that does one thing better than a general-purpose GPU, in exchange for years of tape-out risk and a captive supply chain. OpenAI is a later entrant to that pattern, and China’s domestic accelerator makers are running a parallel version of the same bet, forced by export controls rather than chosen for economics. The common thread across all of them: when a single customer’s workload is big enough, custom silicon beats a general-purpose part on cost per unit of useful work, even after absorbing the design cost.
Why OpenAI cares about latency specifically
Throughput per watt is a data-center-operator’s number. Latency is a product number, and it is the more interesting one given what OpenAI actually ships now. A chatbot response feels instant if it starts in 300 milliseconds. An agent does not get that luxury: it plans a step, calls a tool, reads the result, and plans the next step, sometimes a dozen times before it produces anything the user sees. Add 200 milliseconds of unnecessary latency to a single-turn chat and a user barely notices. Add it to every step of a fifteen-step coding agent run and the task takes three seconds longer, then thirty, then a full minute once the agent starts retrying failed steps. Latency does not add across an agent’s steps. It compounds.
That is the quiet reason a chip company inside a model company chose to optimize latency this hard, and it shows up in the software story as much as the silicon. Jalapeño’s kernels are written in Gluon, a kernel language built on Triton, and some of the most delicate ones, including the DeepSeek MLA attention kernel, were written by Codex, OpenAI’s own coding agent, without an engineer intervening. Someone on the team also got Doom running on Jalapeño using nothing but Codex prompts, which is a party trick, but the MLA kernel is not: a model company used its own agent to write low-level kernel code for the chip meant to serve that same agent faster. That loop, models optimizing the silicon that serves models, is the more durable story here than any single latency number.
The bottom line is that the results show a very, very significant performance advance over state of the art.
What the benchmark does not show
None of this is independently verified in the way SemiAnalysis’s own benchmarking usually is, and SemiAnalysis says so directly. That is worth taking at face value rather than smoothing over, because the caveats are specific enough to matter.
Jalapeño also cannot train anything. It is an inference-only part, and OpenAI’s own plan is to deploy it in “very small volumes” inside its own datacenters by the end of 2026, with a real production ramp through 2027. Nvidia is not losing an order here. It is losing a monopoly on OpenAI’s inference fleet, which is a different and smaller thing than the headlines about Jalapeño “beating Blackwell” imply.
What this means for infrastructure spending
Treat Jalapeño as what it is: proof that a full-stack model company can now design and ship its own inference silicon inside eighteen months of a standing start, not proof that Nvidia’s equity story is at risk this year. OpenAI is still buying Nvidia GPUs at a scale that dwarfs anything Jalapeño will touch through 2027, and the company has publicly committed to deploying Nvidia’s Vera Rubin platform at gigawatt scale in the same window Jalapeño is shipping in “very small volumes.” The two are not competing for the same near-term dollars.
What changes is the negotiating position and the long-run cost curve. Every hyperscaler that successfully stands up its own silicon, Google with the TPU, Amazon with Trainium, and now OpenAI with Jalapeño, removes one more large buyer from the pool of customers who have no alternative to Nvidia’s list price. None of them replace Nvidia outright. All of them make the case, watt by watt, that the largest AI labs will keep splitting their inference fleets between merchant GPUs and purpose-built silicon, and that perf-per-watt, not raw FLOPs, is the number that decides how that split moves over the next few years.
Frequently asked questions
- Is OpenAI replacing Nvidia GPUs with Jalapeño?
- No. OpenAI plans very small-volume Jalapeño deployment inside its own datacenters by the end of 2026, with a larger ramp through 2027, while remaining committed to Nvidia GPUs, including the upcoming Vera Rubin platform, at gigawatt scale over the same period. Jalapeño cannot train models, so it only ever addresses part of OpenAI’s compute needs.
- What models were used to benchmark Jalapeño?
- OpenAI and SemiAnalysis tested three open-weight models on the InferenceX benchmark suite: GPT-OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T, run against Nvidia GB200 and GB300 rack systems.
- Who built the Jalapeño chip?
- Jalapeño was co-designed by OpenAI and Broadcom. It is manufactured on TSMC’s N3P process and pairs each compute die with 15.4 TB/s of HBM4 memory bandwidth across 216 GiB per package.
Sources
- OpenAI (2026). Jalapeño’s first results show industry-leading speed and efficiency in AI inference. OpenAI. https://openai.com/index/jalapeno-first-results/
- Shan, B., Xie, M., Nanos, J., et al. (2026). OpenAI Jalapeño: Better Than Nvidia Blackwell. SemiAnalysis (newsletter). https://newsletter.semianalysis.com/p/openai-jalapeno-better-than-nvidia
- Brandom, R. (2026). OpenAI’s Jalapeño chip is built for fast inference at scale, benchmarks show. TechCrunch. https://techcrunch.com/2026/08/25/openais-jalapeno-chip-is-built-for-fast-inference-at-scale-benchmarks-show/