Viral Wire

OpenAI and Broadcom unveil Jalapeño, first custom LLM inference chip

Designed from scratch for LLMs, early tests show strong performance per watt.

Deep Dive

OpenAI and Broadcom unveiled Jalapeño, OpenAI's first custom Intelligence Processor, an LLM inference accelerator built from the ground up. Early testing with GPT‑5.3‑Codex‑Spark shows it will deliver performance per watt substantially better than current state-of-the-art. The architecture reduces data movement and balances compute, memory, and networking to achieve realized utilization much closer to theoretical peak performance. It's the first in a multi-generation platform with Broadcom, enabling deployment of gigawatt-scale data centers with Microsoft and other partners beginning in 2026.

Key Points
  • Jalapeño is OpenAI's first custom AI inference chip, built from scratch with Broadcom for LLM workloads.
  • Early testing with GPT-5.3-Codex-Spark shows substantially better performance per watt than current state-of-the-art.
  • Part of a multi-generation roadmap targeting gigawatt-scale data centers with Microsoft starting in 2026.

Why It Matters

OpenAI’s custom chip could slash inference costs and energy use, accelerating AI deployment at scale.

📬 Get the top 10 AI stories daily