OpenAI develops 'Jalapeño' chip to challenge Nvidia's inference dominance
Custom silicon outperforms Nvidia's Blackwell systems in key efficiency metrics, signaling a strategic shift toward proprietary hardware.
OpenAI has developed a custom AI inference chip codenamed 'Jalapeño' to reduce its dependence on Nvidia hardware. The move marks a strategic pivot toward proprietary silicon to optimize the massive operational costs of running large-scale AI models.
Developed in partnership with Broadcom, the Jalapeño chip is specifically engineered for AI inference. According to internal and benchmark tests, including data from SemiAnalysis' InferenceX, the chip outperforms Nvidia's Blackwell systems in critical inference-efficiency metrics. Specifically, Jalapeño demonstrated superior performance in tokens per user and throughput per kilowatt, suggesting a more energy-efficient way to deliver AI responses at scale.
The push for custom silicon
For years, the AI industry has relied almost exclusively on Nvidia's H100 and Blackwell GPUs for both training and inference. However, the extreme cost of these systems and persistent supply chain bottlenecks have pushed major AI labs to seek alternatives. OpenAI joins a growing list of tech giants, including Google with its Tensor Processing Units (TPUs) and Amazon with its Trainium and Inferentia lines, that are designing specialized hardware tailored to their specific model architectures.
Industry implications
If OpenAI successfully transitions a significant portion of its inference workload to its own silicon, it could fundamentally alter the economics of AI services. By lowering the cost-per-query, OpenAI can improve its margins and potentially lower prices for users. This shift represents a broader industry transition from general-purpose AI accelerators toward highly specialized, proprietary hardware designed for specific workloads.
What to watch
While the efficiency gains are confirmed in benchmarks, the scale of the rollout remains the primary variable. The industry is now watching to see how quickly OpenAI can integrate Jalapeño into its production environment and whether other hyperscalers will accelerate their own silicon programs in response. It remains to be seen if this transition will lead to a permanent reduction in the pricing power of general-purpose GPU providers.