# OpenAI’s Jalapeño chip posts its first public inference results ahead of 2026 deployment

OpenAI’s first custom inference accelerator has moved from announcement to measured hardware: Jalapeño shows strong InferenceX latency and throughput-per-watt results across three public models, with production deployment planned by year-end.

Jalapeño is now working first-party silicon rather than a roadmap item. OpenAI reports materially better latency and throughput per kilowatt than compared Blackwell systems across GPT-OSS, DeepSeek and Kimi workloads, while SemiAnalysis says it inspected the chip and benchmarked it with its open InferenceX suite.

- Status: Active
- Published: 2026-08-26T06:11:55+12:00
- Updated: 2026-08-26T06:11:55+12:00
- Categories: Artificial Intelligence, Cloud & Infrastructure, Compute & AI Infrastructure, Inference & APIs
- Tags: AI accelerators, inference infrastructure, Jalapeño, OpenAI
- Canonical HTML: https://beyondthe.news/dossiers/openai-jalapeno-inference-chip-first-benchmarks-deployment

## What changed

OpenAI published the first measured results for Jalapeño, its custom LLM inference accelerator developed with Broadcom. On public InferenceX workloads using GPT-OSS 120B, DeepSeek R1 670B and Kimi K2.5 1T, OpenAI reports 1.5–1.9× higher peak throughput per kilowatt and 1.7–3.6× lower end-to-end latency than the comparison systems at the tested operating points. SemiAnalysis, which runs InferenceX, separately says OpenAI invited its team to inspect the hardware and benchmark it with the suite. OpenAI plans to begin deploying Jalapeño inside its own compute infrastructure by the end of 2026 while continuing to use accelerators from NVIDIA and other suppliers.

## Why it matters

The important development is not merely that OpenAI designed a chip; it now has functioning first-generation silicon with externally inspectable benchmark methodology and a deployment timetable. If the measured efficiency survives production qualification, OpenAI can serve more tokens from a fixed power envelope and reduce latency for highly sequential agent workloads, improving capacity economics and potentially the cost/speed envelope exposed through its API. It also adds another serious custom-accelerator path to an inference market dominated by merchant GPUs. Builders cannot buy Jalapeño directly today, however, and OpenAI has not translated the hardware gains into API pricing or service-level commitments.

## The first-generation chip is now measured across public models

OpenAI tested Jalapeño on GPT-OSS 120B, DeepSeek R1 670B and Kimi K2.5 1T rather than only on proprietary OpenAI models. It reports a better latency/throughput-per-watt frontier across the tested operating range, including roughly 1.9× higher peak mixed-token throughput per kilowatt on GPT-OSS and 1.7× on DeepSeek R1. The exact advantage changes with workload and operating point, so no single multiplier describes the chip.

## InferenceX makes the comparison more inspectable than a private vendor benchmark

InferenceX is maintained by SemiAnalysis as an open, reproducible inference benchmark with public recipes, hardware runs and artifacts. SemiAnalysis says OpenAI invited its team to inspect Jalapeño and benchmark the chip with InferenceX. That is stronger evidence than an internal slide deck, although Jalapeño results still depend on the selected models, precisions, server configurations and power-normalization assumptions, and production deployments may behave differently.

## Power efficiency is central to the economics

OpenAI normalizes throughput using published package power ratings: Jalapeño is rated at 700 W and OpenAI says sustained measured power stayed at or below 550 W on the tested workloads. The company reports 1.5–1.9× more AI work per watt at peak throughput across the three public models. For a provider constrained by datacenter power, successful production results would translate into more served work from the same electrical capacity rather than merely faster individual responses.

## Low latency matters disproportionately for agents

OpenAI reports 1.7–3.6× lower end-to-end latency on the public comparisons and 2.1–4.1× higher performance for highly interactive operating points. Sequential agent workflows compound inference delay across many calls, so reducing per-turn latency can shorten an entire task even when model quality is unchanged. The builder-facing significance will depend on whether these gains appear in actual OpenAI API tiers and at what price.

## Production qualification is still unfinished

OpenAI says it is continuing production qualification, software maturation and validation across additional models, with deployment inside its infrastructure planned by the end of 2026. Jalapeño therefore remains an infrastructure transition rather than a currently purchasable product. Gen 2 is already in development and Gen 3 is being designed, indicating that OpenAI sees custom inference silicon as a continuing platform rather than a one-off experiment.

## Key details

- OpenAI announced first measured Jalapeño results on August 25, 2026.
- The public comparisons cover GPT-OSS 120B, DeepSeek R1 670B and Kimi K2.5 1T on InferenceX.
- OpenAI reports 1.5–1.9× higher peak throughput per kilowatt across those workloads than the comparison systems.
- It reports 1.7–3.6× lower end-to-end latency across the three public model comparisons.
- Jalapeño is rated at 700 W; OpenAI says sustained measured power was at or below 550 W on the tested workloads.
- SemiAnalysis says its team inspected the chip and benchmarked it using its InferenceX suite.
- OpenAI plans to begin deploying Jalapeño in its compute infrastructure by the end of 2026.
- The company says it will continue using NVIDIA and other partner accelerators alongside its custom silicon.

## Builder takeaways

- Do not convert Jalapeño benchmark multipliers directly into expected API speed or cost; OpenAI has not yet exposed a Jalapeño-specific service tier or price.
- For latency-sensitive agent products, watch end-to-end task latency rather than output tokens per second alone, because sequential calls amplify serving delays.
- If your business depends heavily on OpenAI inference pricing, model custom silicon as a potential medium-term cost/capacity lever but keep current budgets anchored to published API prices.
- When comparing accelerator claims, use the full latency-throughput curve, model/precision configuration and power basis rather than a single peak tokens-per-second number.
- Infrastructure vendors should treat Jalapeño as evidence that first-party model providers are vertically integrating serving hardware, which can change demand for merchant accelerators and hosted inference stacks.

## What to watch

- Whether Jalapeño enters production on schedule by the end of 2026 and which OpenAI products or API tiers use it first.
- Any API pricing, rate-limit or latency changes explicitly attributed to Jalapeño deployment.
- Independent or public InferenceX runs after production software and system configurations stabilize.
- Reliability, yield, networking and fleet-management evidence at datacenter scale.
- How Gen 2 and Gen 3 change memory capacity, interconnect, precision support and workload scope.
- Whether OpenAI makes Jalapeño capacity available to external customers in any direct form.

## Uncertainties

- Most performance claims are published by OpenAI, although SemiAnalysis says it inspected and benchmarked the chip using its suite.
- Benchmark results depend on model, precision, sequence lengths, serving software, system topology and the power-normalization method; they are not universal workload multipliers.
- Jalapeño has not yet completed production qualification or fleet-scale deployment.
- OpenAI has not disclosed Jalapeño manufacturing economics or translated the measured efficiency into customer pricing.
- OpenAI says internal frontier-model gains are larger, but those results are not independently inspectable because the models are proprietary.

## Sources

- [Jalapeño’s first results show industry-leading speed and efficiency in AI inference](https://openai.com/index/jalapeno-first-results/) — OpenAI · primary/vendor · 2026-08-25T00:00:00+12:00. Primary source for benchmark results, measurement approach, power figures and deployment plan.
- [OpenAI’ Jalapeño: Better Than Nvidia Blackwell](https://newsletter.semianalysis.com/p/openai-jalapeno-better-than-nvidia) — SemiAnalysis · specialist/independent · 2026-08-25T00:00:00+12:00. Independent specialist account stating SemiAnalysis inspected the chip and benchmarked it with InferenceX; detailed article is partly paywalled.
- [About InferenceX](https://inferencex.semianalysis.com/about) — InferenceX / SemiAnalysis · benchmark methodology. Explains the open benchmark’s reproducibility, public recipes, run artifacts, metrics and limitations.
- [OpenAI and Broadcom unveil LLM-optimized inference chip](https://openai.com/index/openai-broadcom-jalapeno-inference-chip/) — OpenAI · primary/vendor · 2026-06-24T00:00:00+12:00. Background on the Broadcom partnership and first-generation accelerator program.

