Connect with us

NEWS

OpenAI’s Jalapeño Chip Outruns Nvidia While Still Needing It

OpenAI says Jalapeño does more AI work per watt than Nvidia GB300, eight days after Nvidia guaranteed $105 billion of Ohio leases.

OpenAI said Tuesday that Jalapeño, its first custom inference chip, did 1.5 to 1.9 times more AI work per watt than Nvidia’s GB200 and GB300 systems. End-to-end latency fell 1.7 to 3.6 times across three public models, the company wrote. Those figures landed eight days after Nvidia agreed to guarantee up to $105 billion of OpenAI’s Ohio data-center leases, a campus where Nvidia will be the exclusive chip supplier.

The ASIC, built with Broadcom and shown at Hot Chips, is still slated for only small volumes inside OpenAI by year-end. Hardware vice president Richard Ho told Bloomberg after the briefing that Nvidia remains a really good partner and that the lab still needs a lot of Nvidia silicon.

Jalapeño Posts 1.9 Times the Work per Kilowatt

OpenAI’s engineering blog put the chip on SemiAnalysis’s InferenceX suite and compared it with the best published GB200 and GB300 points at the time. The headline band is 1.5 to 1.9 times more work per watt at peak throughput and 1.7 to 3.6 times lower end-to-end latency, matching the 1.5 to 1.9 times more work per watt range in the company’s own tables. For highly interactive loads, OpenAI reported 2.1 to 4.1 times higher performance.

The tests used three open models rather than a private ChatGPT snapshot, which is why the appendix can be read in public. GPT-OSS 120B was scored against a GB200 rated at 1,200 watts. DeepSeek R1 670B and Moonshot’s Kimi K2.5 1T were scored against a GB300 rated at 1,400 watts. Jalapeño’s published package rating is 700 watts, and OpenAI said sustained draw stayed at or below 550 watts on these jobs.

Model Comparison chip Peak mixed TPS per kW End-to-end latency
GPT-OSS 120B GB200, 1,200 W 85,448 vs. 44,960 (about 1.9x) 1.03 s vs. 1.80 s (about 1.7x)
DeepSeek R1 670B GB300, 1,400 W 19,641 vs. 11,781 (about 1.7x) 1.65 s vs. 5.99 s (about 3.6x)
Kimi K2.5 1T GB300, 1,400 W 18,195 vs. 11,862 (about 1.5x) 1.56 s vs. 5.31 s (about 3.4x)

The widest gaps sit at the other system’s previous-best time-between-tokens setting, not at peak. On DeepSeek R1, OpenAI logged about 104.3 times more mixed throughput per kilowatt at 169.41 tokens per second per user (12,258 vs. 118). GPT-OSS 120B showed about 53.7 times at 535.28 tok/s/user, and Kimi K2.5 about 56.1 times at 182.46 tok/s/user. Those multipliers describe a matched low-latency point on the Pareto curve, which is a different claim from winning every rack in a data hall.

On GPT-OSS, minimum time between tokens fell to 0.69 ms from 1.87 ms, or 1,459 vs. 535 tokens per second per user. DeepSeek R1’s minimum TBT was 1.43 ms vs. 5.90 ms (700 vs. 169 tok/s/user). Kimi’s was 1.44 ms vs. 5.48 ms (694 vs. 182 tok/s/user). OpenAI said internal runs on its own frontier models widened the lead further, without publishing those traces.

How OpenAI Measured the Chip Against GB200 and GB300

InferenceX, formerly InferenceMAX, is an Apache-2 research platform from SemiAnalysis that reruns popular open serving stacks as software changes. The public InferenceX open-source inference benchmark currently lists GB300 NVL72, GB200 NVL72, several AMD MI parts, and Hopper cards as official hardware. Jalapeño is not on that supported-SKU table. OpenAI used the method and compared against recorded Nvidia points, and Tom’s Hardware reported that SemiAnalysis ran InferenceX with OpenAI engineers in the company’s lab.

Power math is where the slide can move. OpenAI normalized to each accelerator’s published chip rating, 700 W for Jalapeño against 1,200 W and 1,400 W for the Nvidia parts, even though Jalapeño’s measured draw was lower. Tom’s Hardware, covering the Hot Chips talk, said an appendix that uses all-in utility power (1.18 kW for Jalapeño against 2.55 kW for the GB300) narrows the gaps, and that a GB300 running multi-token prediction shrinks the peak efficiency lead to roughly 1.5 times. Vera Rubin, Nvidia’s next platform, was not in the public comparison.

WHAT WE KNOW

  • The yardstick: OpenAI scored three public models on InferenceX at a nominal 8k/1k setting and reported package TDP in the appendix.
  • The watt ratings: Jalapeño is listed at 700 W, the GB200 comparison at 1,200 W, and the GB300 comparisons at 1,400 W.
  • The partners in the room: SemiAnalysis said it ran the suite with OpenAI staff on first-party silicon, and Tom’s Hardware quoted the firm describing the part as beating every Nvidia, AMD, and Google chip it has been able to test.

WHAT IS UNCONFIRMED

  • Independent dashboard listing: Jalapeño still does not appear among InferenceX’s officially supported SKUs, so outsiders cannot yet replay the same run from the public chart.
  • Production power and yield: Sustained 550 W in the lab is not the same as fleet draw, defect rates, or rack-level cooling once volume starts.
  • A Rubin matchup: Nvidia’s newer platform was left out, so the published tables stop at Blackwell-class GB200 and GB300 points.

Some recaps on Wednesday already treated the result as a Rubin win. It is not. The public appendix is GB200 on GPT-OSS and GB300 on the two larger models, using single-token prediction in the main cuts.

Eight Days After a $105 Billion Nvidia Backstop

On August 17, Reuters reported that Nvidia agreed to guarantee up to $105 billion of lease and related obligations for an OpenAI-leased campus at SB Energy’s PORTS-Pike site in Pike County, Ohio, and to invest $1.5 billion in the developer. Axios, working from the same announcement, said Nvidia will be the exclusive provider of chips at that site. The campus is planned for as much as 8 gigawatts of IT load, with Nvidia’s written cap covering the first 4.25 gigawatts and an option on the rest, and Reuters said the first 800 megawatts is expected in 2028 on a 20-year lease.

That is a residual-value backstop, not a $105 billion check. It still ties OpenAI’s largest disclosed U.S. hall to Nvidia iron. Eight days later the same customer walked into Hot Chips with a Broadcom ASIC and a slide that puts GB300 behind on watts and wait time.

  1. June 24, 2026: OpenAI and Broadcom unveil Jalapeño, with Hock Tan and Charlie Kawwas handing a wafer to Sam Altman and Greg Brockman, and promise gigawatt-scale deployment with Microsoft and other partners.
  2. August 17, 2026: Nvidia files and announces credit support capped at $105 billion for OpenAI’s Ohio leases and says it will be the exclusive chip vendor there.
  3. August 25, 2026: OpenAI publishes InferenceX numbers and tells reporters the chip offers lower latency and higher throughput together, with Ho adding that the company still needs a lot of Nvidia.
  4. Late 2026: OpenAI plans to start placing Jalapeño in its own compute fleet in small volumes.
  5. 2027: Volume is supposed to rise, while Cerebras, in a post on Tuesday, said it still expects its next system to sit with Jalapeño in OpenAI’s fast and ultrafast lineup.

Ho’s line to Bloomberg is the cleanest summary of the week: Nvidia is a really good partner, and OpenAI continues to need a lot of Nvidia. The Ohio hall makes that concrete. The InferenceX appendix makes the other half concrete. Both can be true at once, which is why the “Nvidia killer” frame does not hold.

Google, Amazon, Meta and Microsoft Got Here First

Jalapeño is OpenAI’s first named inference processor, not the industry’s. Google has been iterating TPUs for years and now splits training and inference parts. Amazon has Inferentia for serving and Trainium for training. Meta has already put hundreds of thousands of MTIA chips on ranking and ads inference, per coverage of its March roadmap, and has more generations queued with Broadcom. Microsoft’s Maia 200 is an inference part already going into Azure. OpenAI is joining that club with a first public scorecard, while still buying the merchant GPUs everyone else also buys.

  • Google TPUs: Longest-running custom line, now with separate train and inference spins and a Broadcom implementation history that predates Jalapeño.
  • Amazon Inferentia and Trainium: A two-chip split that treats serving cost as its own problem, which is the same bet OpenAI is making.
  • Meta MTIA: Already in volume on internal inference, with later chips aimed at generative loads rather than a first lab demo.
  • Microsoft Maia: An Azure inference accelerator from OpenAI’s closest cloud partner, built while OpenAI was still drawing Nvidia for almost all of its own serving.

A May roundup in Tom’s Hardware put custom ASIC server shipments on track for 27.8 percent of the 2026 market and 44.6 percent annual growth, against 16.1 percent growth for merchant GPUs, with Nvidia still holding most of the revenue. Jalapeño does not flip that mix this year. It does put OpenAI on the same side of the table as the other labs that got tired of paying GPU list price for decode.

Broadcom is the quiet constant. It already helps Google and Meta ship custom AI silicon, and now it is the implementation and networking partner on OpenAI’s first Intelligence Processor, including Tomahawk switch silicon. Celestica is on boards, racks, and system integration. Taiwan’s TSMC is the foundry, according to Reuters’ June report. OpenAI designed the architecture; it did not open a fab.

OpenAI Used Its Own Models to Design the Chip

The June announcement said Jalapeño went from initial design to manufacturing tape-out in nine months, a span OpenAI called the fastest ASIC cycle it knows of in this class of parts. The August blog said earlier models helped design and bring up the chip, while later ones now help program it. For selected GPT-OSS attention and mixture-of-experts blocks, OpenAI said AI-written implementations ran 1.5 to 1.8 times faster than the human-expert versions. Those figures apply to the chosen blocks, not the whole model.

We optimized the architecture around the kernels, memory movement, networking, and serving patterns that matter most for frontier AI models. Based on early testing, Jalapeño will efficiently execute our most important workloads close to the hardware’s theoretical limits.

Richard Ho, OpenAI hardware program lead, June 24, 2026 company announcement

The same June post described a nine-month design to tape-out cycle and a blank-slate inference layout rather than a general GPU cut down for chat. Prefill is compute-heavy. Decode is memory-bandwidth heavy. Agents bounce between the two, so a chip that wins one phase can stall on the other. OpenAI says it kept KV cache local, made the network part of the architecture, and treated the rack as one connected domain so a request does not keep walking across slow links.

Using Codex with GPT-Astra, the hardware team brought three open-weight models that were not in the original production plan up to high performance in two months. One of those, the gpt-oss-120b open-weight reasoning model, is a 117 billion parameter mixture-of-experts part with 5.1 billion active weights, released in August 2025 under Apache 2.0 and small enough, in MXFP4, to fit on a single 80 GB GPU. DeepSeek R1 and Kimi K2.5 were the other two public loads. That spread is the company’s answer to the usual custom-silicon jab, that an ASIC only looks good on the model it was born for.

Tom’s Hardware, working from the Hot Chips package photos and slides, said each Jalapeño ships six HBM4 stacks for 216 GiB at 15.4 TB/s, against 288 GB of HBM3E on a 1,400 W GB300, and that OpenAI’s talk treated exposing aggregate HBM bandwidth, not adding more of it, as the bottleneck. If those memory figures hold in volume, OpenAI becomes another bidder in an HBM4 queue that Samsung, SK hynix, and Micron have already sold through 2027, per the same report. That is a supply story, not a FLOPS story.

When OpenAI Plans to Deploy Jalapeño

OpenAI said it will begin putting Jalapeño into its own compute infrastructure by the end of the year, then raise volume in 2027. It did not give a chip count. Ho, in The Verge’s briefing, said existing systems usually trade latency against throughput and that Jalapeño offers the “best of both worlds.” He also said the company does not expect this part to replace its whole lineup.

The company account followed that video with a note that Gen 2 is deep in development and Gen 3 is taking shape. Tom’s Hardware, citing Bloomberg, said the second chip is approaching tape-out within months. Until those parts exist in halls, ChatGPT speed, Codex steps, and API price still ride on Nvidia, AMD, and the Cerebras systems OpenAI already uses for ultrafast serving. Cerebras spent Tuesday congratulating the Jalapeño team and pointing at a 2027 production mix that includes both.

Work per watt is the constraint that makes a 700 W part worth this much ink. Towns are already slowing new halls, including a Michigan township’s one-year freeze on new data centers, while labs try to pull more tokens from the megawatts they already have. OpenAI’s blog is explicit about the business math: more useful work from the same power can let revenue grow faster than the cost to serve. That is an operating-leverage claim, and it only pays once the chip is in the fleet, not while it is an appendix.

Gen 2 Is Already in Development

The published tables are first silicon on three open models, scored with a public method, on a 700 W package that drew 550 W in the lab. They do not retire GB300, they do not cover training, and they do not land in Ohio, where Nvidia has the exclusive. They do show why OpenAI spent nine months and a Broadcom contract to stop treating decode as a GPU leftover.

Hock Tan told Reuters in June that Jalapeño was in the same class as Nvidia Blackwell and Google TPUs, and he told Bloomberg then that early samples pointed to cost cuts on the order of half versus a typical AI GPU. Those were preview lines. Tuesday’s appendix is the first time outsiders can see the actual mixed TPS per kilowatt. The next test is whether small volumes at year-end behave like the Pareto plots, and whether HBM4 allocations and software maturity keep up with Gen 2’s tape-out.

OpenAI will keep buying accelerators from Nvidia and other partners for training and for inference, the blog said, because demand is still larger than any one platform. Jalapeño is the first generation of that side bet, scored in public, financed on the other side by the same vendor it just put in the slower column.

Click to comment

Leave a Reply

Your email address will not be published. Required fields are marked *