OpenAI's upcoming Jalapeño AI chip outperforms NVIDIA GB300 in inference tests

Last year, OpenAIannounceda major partnership with Broadcom to deploy 10 gigawatts of custom AI accelerators by 2029. In June 2026, OpenAI and Broadcomunveiled Jalapeño, OpenAI's first custom AI inference chip, designed specifically for large language models and agentic workloads.

At the time, OpenAI said Jalapeño would deliver substantially better performance per watt than existing accelerators, but it did not publish detailed numbers. Many in the industry were skeptical of those claims, since this is OpenAI's first attempt at designing its own chip, while established players such as NVIDIA, Google, and Amazon have years of experience developing and deploying multiple generations of high-performance AI chips.

Today, OpenAI released its first performance results for Jalapeño, and they appear impressive on paper. OpenAI tested Jalapeño using SemiAnalysis' public InferenceX benchmark across GPT-OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T. According to the results, Jalapeño delivered 1.5x to 1.9x more AI work per watt at peak throughput and 1.7x to 3.6x lower end-to-end latency than the NVIDIA GB200 and GB300-based comparison systems. For highly interactive workloads, OpenAI claims a 2.1x to 4.1x performance advantage.

For example, while running DeepSeek R1, Jalapeño delivered 19,641 mixed tokens per second per kilowatt, compared to 11,781 for the NVIDIA GB300. End-to-end latency was also reduced from 5.99 seconds to 1.65 seconds. With Kimi K2.5, Jalapeño achieved about 1.5x higher peak performance per watt and 3.4x lower latency.

Jalapeño also has a rated power consumption of just 700W, though OpenAI says sustained power stayed at or below 550W during its tests. For comparison, a GB300 system has a rated power consumption of 1400W. OpenAI highlighted that the chip was designed alongside its memory, networking, software, and rack-scale systems to minimize data movement.

Once again, OpenAIhighlightedthat its AI models helped design and program the chip. The Jalapeño chip went from initial design to tapeout in just nine months. OpenAI plans to begin deploying Jalapeño in its infrastructure by the end of 2026. A Gen 2 chip is already deep in development, while work on Gen 3 has also started.

Despite the promising Jalapeño results and future custom silicon roadmap, OpenAI says it will continue deploying NVIDIA and other AI accelerators for both training and inference to meet growing demand.

Postar um comentário

Postagem Anterior Próxima Postagem