OpenAI used its first-ever Hot Chips appearance to put numbers on the slide for its in-house inference ASIC, Jalapeño, and to say the part is already serving Codex. The talk ran late on Tuesday 25 August at Stanford. It is the first time the company has shown its own silicon in public with figures attached.

What was actually disclosed

Per chip: 13.4 PFLOP/s of MXFP4-by-MXFP4 matrix compute, 15.4 TB/s of HBM4 bandwidth across 216 GiB, inside a 700-watt package. The interconnect is a half-flattened two-level Clos built on Broadcom Tomahawk 6 switches — a 600 GB/s local domain across a 128-ASIC cluster, a 200 GB/s global domain across the full system. Broadcom and Celestica were acknowledged as partners.

The schedule is the impressive part

OpenAI gave the timeline verbatim: architecture concept in late 2024, RTL freeze in 2025, tape-out in late 2025, Codex running on the chip in early 2026, with ChatGPT to follow. Roughly two years from a blank page to serving a production workload is fast for a first-pass accelerator.

What the common framing gets wrong

Three errors, and they compound. First, 27 EFLOP/s is a system number — 13.4 PFLOP/s times 2,048 chips, alongside 432 TiB of aggregate memory. It will be quoted next to Nvidia per-GPU figures, where it means nothing. Second, MXFP4 is the smallest and most favourable precision available. Four-bit microscaled matrix math is not comparable to a dense FP8 or BF16 number, and OpenAI published no dense figure at all, so any "beats Nvidia" arithmetic that crosses precisions is void. Third, the 700 W versus GB200's 1.2 kW and GB300's 1.4 kW is a package-power comparison, not a performance-per-watt result — those Nvidia parts are CPU-plus-GPU superchips carrying training capability, while Jalapeño is inference-only, and no shared benchmark was shown.

And it is first silicon

The software stack was brought up between A0 returning to the lab and the talk. That is a lab-to-limited-service story. Nothing about process node, wafer volume, yield or fleet size was disclosed — which is exactly the set of facts that separates a working chip from a supply line.