What changed
OpenAI published the first measured results for Jalapeño, its custom LLM inference accelerator developed with Broadcom. On public InferenceX workloads using GPT-OSS 120B, DeepSeek R1 670B and Kimi K2.5 1T, OpenAI reports 1.5–1.9× higher peak throughput per kilowatt and 1.7–3.6× lower end-to-end latency than the comparison systems at the tested operating points. SemiAnalysis, which runs InferenceX, separately says OpenAI invited its team to inspect the hardware and benchmark it with the suite. OpenAI plans to begin deploying Jalapeño inside its own compute infrastructure by the end of 2026 while continuing to use accelerators from NVIDIA and other suppliers.
Why it matters
The important development is not merely that OpenAI designed a chip; it now has functioning first-generation silicon with externally inspectable benchmark methodology and a deployment timetable. If the measured efficiency survives production qualification, OpenAI can serve more tokens from a fixed power envelope and reduce latency for highly sequential agent workloads, improving capacity economics and potentially the cost/speed envelope exposed through its API. It also adds another serious custom-accelerator path to an inference market dominated by merchant GPUs. Builders cannot buy Jalapeño directly today, however, and OpenAI has not translated the hardware gains into API pricing or service-level commitments.
The first-generation chip is now measured across public models
OpenAI tested Jalapeño on GPT-OSS 120B, DeepSeek R1 670B and Kimi K2.5 1T rather than only on proprietary OpenAI models. It reports a better latency/throughput-per-watt frontier across the tested operating range, including roughly 1.9× higher peak mixed-token throughput per kilowatt on GPT-OSS and 1.7× on DeepSeek R1. The exact advantage changes with workload and operating point, so no single multiplier describes the chip.
InferenceX makes the comparison more inspectable than a private vendor benchmark
InferenceX is maintained by SemiAnalysis as an open, reproducible inference benchmark with public recipes, hardware runs and artifacts. SemiAnalysis says OpenAI invited its team to inspect Jalapeño and benchmark the chip with InferenceX. That is stronger evidence than an internal slide deck, although Jalapeño results still depend on the selected models, precisions, server configurations and power-normalization assumptions, and production deployments may behave differently.
Power efficiency is central to the economics
OpenAI normalizes throughput using published package power ratings: Jalapeño is rated at 700 W and OpenAI says sustained measured power stayed at or below 550 W on the tested workloads. The company reports 1.5–1.9× more AI work per watt at peak throughput across the three public models. For a provider constrained by datacenter power, successful production results would translate into more served work from the same electrical capacity rather than merely faster individual responses.
Low latency matters disproportionately for agents
OpenAI reports 1.7–3.6× lower end-to-end latency on the public comparisons and 2.1–4.1× higher performance for highly interactive operating points. Sequential agent workflows compound inference delay across many calls, so reducing per-turn latency can shorten an entire task even when model quality is unchanged. The builder-facing significance will depend on whether these gains appear in actual OpenAI API tiers and at what price.
Production qualification is still unfinished
OpenAI says it is continuing production qualification, software maturation and validation across additional models, with deployment inside its infrastructure planned by the end of 2026. Jalapeño therefore remains an infrastructure transition rather than a currently purchasable product. Gen 2 is already in development and Gen 3 is being designed, indicating that OpenAI sees custom inference silicon as a continuing platform rather than a one-off experiment.