OpenAI’s Jalapeño Chip Tops Blackwell
OpenAI’s first custom inference chip reportedly delivers 1.5–1.9× more performance per watt and 1.7–3.6× lower latency than Nvidia GB200 and GB300 systems across GPT-OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T. The results come from SemiAnalysis’s InferenceX benchmark and OpenAI testing.
Jalapeño looks like a surprisingly strong first-generation inference ASIC, but the headline numbers need qualification: the chip is still in limited deployment, and the comparisons are not yet independently reproducible at scale.
- –Its strongest advantage is serving efficiency, directly targeting OpenAI’s enormous inference costs.
- –Generalized performance across OpenAI, DeepSeek, and Moonshot models makes this more credible than a narrowly optimized accelerator.
- –Lower latency matters especially for agentic workloads, where delays compound across many sequential model calls.
- –Jalapeño is inference-only, so OpenAI remains dependent on Nvidia and other partners for training capacity.
- –The bigger strategic shift is full-stack control: OpenAI can co-design models, kernels, serving software, memory, networking, and silicon.
DISCOVERED
2h ago
2026-08-25
PUBLISHED
2h ago
2026-08-25
RELEVANCE
AUTHOR
mark_k