Developer Tools

OpenAI unveils Jalapeño, its first custom AI inference chip with Broadcom

Early tests show significantly better performance-per-watt than current state-of-the-art alternatives.

Deep Dive

OpenAI has officially unveiled its first custom-designed inference processor, Jalapeño, built in collaboration with Broadcom. The chip is purpose-built for running pre-trained AI models in response to user commands, a process known as inference. According to OpenAI, early testing shows Jalapeño achieving significantly better performance-per-watt than current state-of-the-art alternatives, with a particular emphasis on low operating costs for real-time coding models. OpenAI president Greg Brockman explained that the company identified underserved workloads and designed the chip to accelerate what's possible, leveraging its own AI models in the development process.

This move is part of OpenAI's broader strategy to reduce dependence on Nvidia GPUs, following similar custom chip efforts by Google and Amazon. While performance-intensive tasks like pre-training will likely still rely on Nvidia hardware, even small reductions in inference costs could significantly improve OpenAI's bottom line. The company emphasized that it operates across the entire stack—from chip architecture and kernels to networking, scheduling, and product experience—allowing each layer to be optimized for faster, more reliable, and more affordable models. With agentic products like Codex already in development, Jalapeño represents a key infrastructure step in OpenAI's vertical integration play.

Key Points
  • Jalapeño is OpenAI's first custom inference chip, designed and manufactured with Broadcom.
  • Early results show significantly better performance-per-watt than current state-of-the-art alternatives, especially for real-time coding models.
  • The chip is part of OpenAI's strategy to reduce reliance on Nvidia GPUs, while pre-training will still likely use Nvidia hardware.

Why It Matters

Custom chips like Jalapeño could slash inference costs, making AI agents and coding tools more affordable at scale.

📬 Get the top 10 AI stories daily