Anime, manga, and games, with a take · A Yukimedia publication

← all stories other 1 sources · 1h ago ·

OpenAI Publishes Benchmark Results for Custom Inference Chip Jalapeno

Jalapeño's published benchmarks show it overcoming the usual throughput-versus-latency trade-off in a single architecture, with the advantage growing on larger and more demanding workloads.

Reporting from 1 source: GIGAZINE.

OpenAI Publishes Benchmark Results for Custom Inference Chip Jalapeno

OpenAI held a briefing on Jalapeño, a custom inference chip co-developed with Broadcom, and published benchmark results. The chip reportedly achieves both high throughput and low latency in a single architecture, a combination that typically requires a trade-off. OpenAI plans to deploy the first generation within its internal computing infrastructure by the end of the year.

Jalapeño was announced in June 2026 as a custom inference chip optimized for large language models. Since then, OpenAI has run repeated tests on the chip and the systems built around it.

Across three models, GPT OSS 120B, DeepSeek R1, and Kimi K2.5 1T, peak throughput improved AI processing per 1kW by 1.5 to 1.9 times and cut end-to-end latency by 1.7 to 3.6 times. For highly interactive workloads, the improvement rose to 2.1 to 4.1 times. Peak decode throughput improved across all three models, and throughput per 1kW exceeded previous records.

OpenAI plans to deploy the first generation inside its own infrastructure by the end of the year. A second generation is in development, and a third generation is taking shape.

Synthesized by Yomimono from the 1 cited source below, including Japanese-language reporting where cited, then editorially reviewed before publishing.

Sources