Nine days after Nvidia agreed to backstop up to $105 billion in financing for an OpenAI-leased data center campus in Ohio, OpenAI walked onstage at Hot Chips in Santa Clara and told the room its own silicon beats Nvidia's. On Tuesday, August 25, the company published the first benchmark results for Jalapeño, the inference ASIC it co-designed with Broadcom, claiming 1.5x to 1.9x more throughput per kilowatt and 1.7x to 3.6x lower end-to-end latency than Nvidia's GB200 NVL72 and GB300 NVL72 rack systems. For highly interactive workloads, OpenAI put the gap at 2.1x to 4.1x. The numbers ricocheted through the market a day before Nvidia's earnings call.
Jalapeño is a 700W part going up against accelerators rated at 1,200W and 1,400W. Each package pairs its compute die with six HBM4 stacks totaling 216 GiB at 15.4 TB/s; the GB300 carries 288GB of slower HBM3E at 1,400W, so per watt of rated power OpenAI's chip carries roughly 50% more memory. The tests ran on SemiAnalysis's public InferenceX suite across three open models — GPT-OSS 120B, DeepSeek R1 670B, and Moonshot's 1-trillion-parameter Kimi K2.5 — with the widest margins at low-latency operating points, where OpenAI claims 8.6x to 104.3x more throughput per kilowatt at the GB300's fastest previous time-between-tokens settings.
These are vendor-supplied results. OpenAI ran the benchmarks and produced the numbers; SemiAnalysis says it executed InferenceX alongside OpenAI engineers in OpenAI's own lab and verified some runs on-site. That is more third-party contact than most first-party chip claims get, and less than independent testing. SemiAnalysis was unequivocal in its writeup, describing the part as "beating every Nvidia, AMD, and Google chip we have been able to test." Its CEO, Dylan Patel, posted that "Usually first generation chips aren't competitive, but OpenAI is beating Nvidia Blackwell and even Rubin."
What Work Per Watt Does and Does Not Measure
Throughput per kilowatt is tokens generated per unit of electrical power. It is the right metric for an industry whose binding constraint has shifted from wafers to substations, and it is where a purpose-built inference ASIC should beat a general-purpose GPU. But the normalization choices matter, and OpenAI made several that cut in different directions.
OpenAI normalized to each accelerator's published package TDP rather than measured draw, even though it says Jalapeño's sustained power stayed at or below 550W in testing — a choice that understates its own advantage. An appendix using all-in utility power per accelerator, 1.18 kW for Jalapeño against 2.55 kW for the GB300, produces narrower gaps. So does pitting Jalapeño against a GB300 running multi-token prediction, where the peak efficiency lead shrinks to roughly 1.5x. The headline comparisons ran Jalapeño's single-token prediction against GB300 configurations doing the same, and excluded speculative decoding, even though production Nvidia deployments commonly use those optimizations.
Then there is what the metric ignores entirely. Perf/watt says nothing about capital cost per token, nothing about total cost of ownership — SemiAnalysis puts Jalapeño and Vera Rubin roughly even there — and nothing about training, the workload where Nvidia remains unchallenged and where Jalapeño does not compete at all. Crucially, Jalapeño was not benchmarked against Vera Rubin, the platform slated to power the first gigawatt of Nvidia systems OpenAI itself agreed to deploy in the second half of 2026. Nvidia and AMD have already published results on larger models, including DeepSeek V4 Pro and Kimi K3, that Jalapeño has not run.
Analysts split accordingly. Adrien Sanchez of Yole Group told CNBC the results show a "hyperscaler-designed chip can now match or beat Nvidia's Blackwell-class GPUs on inference efficiency," calling it a threat to Nvidia's inference margins. Omdia senior principal analyst Alexander Harrowell called Jalapeño an "impressive achievement, most of all in terms of efficiency." TrendForce's Fion Chiu was more measured: the chip cuts reliance on Nvidia for inference, but for compute-intensive training, "Nvidia GPUs will remain important."
Why It Matters
The engineering timeline may be the more consequential disclosure. Design work began in mid-2024 and the final design went to fabrication in November 2025 — about 16 months end to end, with OpenAI saying only nine months elapsed from first chip design to tapeout-ready blueprint. OpenAI used its own models throughout: older generations assisting chip design, newer ones accelerating the software bring-up. SemiAnalysis drew the obvious conclusion, writing that "The CUDA moat is potentially dead given how fast OpenAI can bring up new models on their silicon." If a lab can stand up a competitive inference stack on new silicon in months rather than years, the software lock-in that has protected Nvidia's pricing power for a decade is a depreciating asset.
The economics remain tangled. Jalapeño is built on a TSMC 3nm-class process, keeping OpenAI queued for the same wafers, HBM allocation and advanced packaging as Blackwell and Rubin. Scaling it across the 10-gigawatt agreement OpenAI signed with Broadcom last October would make the company a major new claimant on HBM4 supply Nvidia dominates through multi-year deals with SK hynix — in a market where Micron told the same conference on August 23 that HBM consumes roughly three times the wafer area of DDR5 for equivalent capacity, and where Nvidia CFO Colette Kress warned Wednesday that memory prices will push gross margin down to 71% to 72% by the fourth quarter of fiscal 2027. Richard Ho, OpenAI's vice president of hardware, put the relationship plainly to Bloomberg after the Nvidia financing deal: "Nvidia is a really good partner, and we continue to need a lot of Nvidia."
Watch three things. First, whether Jalapeño moves past engineering samples: OpenAI plans small-volume deployment in its own data centers by the end of 2026 with a broader rollout in 2027, while Rubin systems are already shipping to customers. Second, whether anyone runs InferenceX on Jalapeño against Vera Rubin without OpenAI in the room. Third, the second-generation part, which Bloomberg reports is approaching tapeout within months, with concept work on a third already underway. First-generation chips are supposed to be uncompetitive. The claim being made here is that the second one will not be a first draft.
“Usually first generation chips aren't competitive, but OpenAI is beating Nvidia Blackwell and even Rubin.”— Dylan Patel, CEO, SemiAnalysis