TechNewsReel
Live

OpenAI's Jalapeño ASIC Outperforms Nvidia Flagships in Power Efficiency

New benchmarks show OpenAI's first in-house inference chip delivers higher throughput and lower latency than Nvidia's GB200 and GB300 systems.

TechNewsReel Newsroom · August 25, 2026

OpenAI has unveiled benchmarks for its first in-house inference ASIC, "Jalapeño," claiming the chip outperforms Nvidia's latest rack systems in both speed and power efficiency. The announcement, made at the Hot Chips conference, marks a significant step in the company's effort to vertically integrate its hardware stack.

Developed in collaboration with Broadcom, the Jalapeño chip is rated at 700W, with measured sustained power levels at or below 550W. This is substantially lower than Nvidia's flagship accelerators, which are rated at 1,200W and 1,400W. According to data presented by OpenAI, Jalapeño delivered 1.5x to 1.9x more throughput per kilowatt and 1.7x to 3.6x lower end-to-end latency than Nvidia GB200 and GB300 systems. These results were derived using the SemiAnalysis InferenceX suite, testing models including Kimi K2.5, GPT-OSS 120B, and DeepSeek R1 670B. On the hardware side, the ASIC features 216 GiB of High Bandwidth Memory (HBM) with a bandwidth of 15.4 TB/s.

The Push for Hardware Independence

This move comes as OpenAI seeks to diversify its hardware strategy and reduce its total reliance on Nvidia. While the company remains deeply entwined with the GPU giant—with Richard Ho, OpenAI VP of Hardware, stating that "Nvidia is a really good partner, and we continue to need a lot of Nvidia"—the development of Jalapeño targets the specific demands of inference. While Nvidia continues to dominate the training workloads required to build large language models, the operational phase of running those models (inference) is where power constraints and costs become most acute.

Implications for AI Scaling

If OpenAI can successfully deploy custom ASICs that are nearly twice as power-efficient as industry-standard flagships, it drastically reduces the operational costs and energy requirements of scaling LLM inference. This shift signals a broader industry trend where AI labs design silicon tailored specifically to their own model architectures rather than relying on general-purpose GPUs. By optimizing the hardware for the specific mathematical operations of its models, OpenAI can bypass some of the overhead inherent in more flexible, general-purpose chips.

Deployment and Future Outlook

OpenAI plans to deploy the Jalapeño chip in its own data centers later this year. The move is also a strategic play amid a severe global HBM shortage, with current capacity reportedly sold through 2027. By designing its own memory specifications and silicon, OpenAI gains more control over its supply chain. Observers will now be watching to see if these conference benchmarks translate to real-world performance at scale and how Nvidia responds to the increasing trend of custom silicon among its largest customers.

Sources

Get a notification when a big story breaks. A few a day at most — no spam.