OpenAI's Jalapeño chip cuts inference costs, rivals Nvidia
Custom ASIC from Broadcom partnership slashes AI serving costs by 2-3x
OpenAI has officially entered the custom silicon race with Jalapeño, an ASIC (application-specific integrated circuit) built exclusively for LLM inference — the process of running trained AI models in real time to respond to user prompts. Unlike Nvidia's GPUs, which are optimized for training, Jalapeño is purpose-built to make inference faster, cheaper, and more power-efficient at scale. The chip was developed in collaboration with Broadcom (which provided silicon implementation and Tomahawk networking) and Celestica (board integration). OpenAI claims the design went from concept to tape-out in just nine months — potentially one of the fastest advanced chip development cycles on record — and that its own AI models helped accelerate parts of the design process.
Early internal testing suggests Jalapeño delivers significantly better performance per watt than current state-of-the-art chips, though OpenAI has not yet released final benchmarks or a full technical report. The chip's architecture reduces data movement and balances compute, memory, and networking — key bottlenecks in AI inference. This move is part of a broader industry trend: major AI companies seeking to reduce dependence on external suppliers like Nvidia. OpenAI president Greg Brockman called it a “full-stack” strategy to improve efficiency and reduce cost. With inference being the daily interaction point for products like ChatGPT and future agent-style services, even small efficiency gains translate into massive cost savings at scale. Jalapeño positions OpenAI alongside Google and Amazon, which have already built their own AI chips, while still relying on Nvidia for training workloads.
- Jalapeño is an ASIC designed exclusively for LLM inference, not training, optimized for speed, cost, and power efficiency.
- The chip was developed in 9 months with AI-assisted design, using Broadcom's Tomahawk networking and Celestica integration.
- Early internal tests show major performance-per-watt improvements over current GPUs, but independent benchmarks are pending.
Why It Matters
Reduces OpenAI's Nvidia dependence and cuts inference costs, enabling cheaper ChatGPT and agent products at scale.