OpenAI and Broadcom unveil Jalapeño, first custom LLM inference chip
Designed from scratch for LLMs, early tests show strong performance per watt.
OpenAI and Broadcom unveiled Jalapeño, OpenAI's first custom Intelligence Processor, an LLM inference accelerator built from the ground up. Early testing with GPT‑5.3‑Codex‑Spark shows it will deliver performance per watt substantially better than current state-of-the-art. The architecture reduces data movement and balances compute, memory, and networking to achieve realized utilization much closer to theoretical peak performance. It's the first in a multi-generation platform with Broadcom, enabling deployment of gigawatt-scale data centers with Microsoft and other partners beginning in 2026.
- Jalapeño is OpenAI's first custom AI inference chip, built from scratch with Broadcom for LLM workloads.
- Early testing with GPT-5.3-Codex-Spark shows substantially better performance per watt than current state-of-the-art.
- Part of a multi-generation roadmap targeting gigawatt-scale data centers with Microsoft starting in 2026.
Why It Matters
OpenAI’s custom chip could slash inference costs and energy use, accelerating AI deployment at scale.