Developer Tools

OpenAI and Broadcom unveil Jalapeño chip for LLM inference at scale

Custom ASIC promises better performance per watt, deployed by year-end.

Deep Dive

OpenAI and Broadcom have introduced Jalapeño, a custom ASIC specifically optimized for large language model inference at data center scale. The chip is the first generation of a long-term project, with development informed directly by OpenAI's researchers and their future model roadmap. According to Broadcom, Jalapeño was designed from the ground up to handle the unique computational patterns of LLMs, which differ from traditional GPU workloads. Early testing indicates performance per watt substantially better than current state-of-the-art, though OpenAI has not yet released full benchmarks—a detailed technical report is expected in the coming months.

The partnership underscores OpenAI's broader strategy to vertically integrate its hardware stack, reducing reliance on suppliers like Nvidia. By controlling both the silicon and the software, the company aims to squeeze more capacity from limited data center resources amid a global compute crunch. Broadcom, already a major player in custom chips for hyperscalers, has seen growing demand from frontier AI labs. The companies confirmed Jalapeño chips will be deployed in data centers before the end of this year, marking a significant step toward purpose-built AI infrastructure.

Key Points
  • Jalapeño is an ASIC designed specifically for LLM inference, not general-purpose computing.
  • Developed in nine months based on OpenAI's internal model and product roadmap.
  • Claims substantially better performance per watt than current options; full benchmarks pending.

Why It Matters

Custom silicon for LLM inference could lower costs, reduce Nvidia dependency, and scale AI deployment.

📬 Get the top 10 AI stories daily