OpenAI's custom inference chip, Jalapeño, has moved from confirmation to measured silicon. In a Q&A with Dr. Ian Cutress of More Than Moore, OpenAI VP of Hardware Richard Ho detailed the part the company is building with Broadcom as design partner and Celestica handling board and rack integration.
The interview follows Ho's HotChips 2026 presentation, where he showed the first measured performance from working Jalapeño silicon alongside Ravi Narayanaswami and Chris Leary. The headline claim: OpenAI's hardware is faster and more efficient than NVIDIA's GB200 and GB300 at inference, and the company used its own AI models to help design it.
OpenAI Jalapeño specs and design
Jalapeño is an inference accelerator, not a training chip. It carries 216 GiB of HBM4 running at 15.4 TB/s, paired with a compute die and an I/O chiplet. Chip power is rated at 700 W peak, with measured sustained draw closer to 550 W.
OpenAI scales the part in two steps: a local domain of 128 accelerators, and a full system of 2,048 accelerators. At four-bit precision, that full system reaches 27 EFLOP/s.
Two design decisions stood out at HotChips. First, OpenAI built a single balanced part for the whole inference workload, while most vendors split prefill, speculative decode and full decode across specialized hardware. Second, OpenAI's own models contributed heavily to the design, from compute work and fitting circuits into the available area to post-optimizing kernels for the hardware.
- Type: inference accelerator (not training)
- Memory: 216 GiB HBM4 at 15.4 TB/s
- Dies: compute die plus I/O chiplet



