Developer Tools

Cerebras and OpenAI launch GPT-5.6 Sol Ultrafast for 11x faster AI

GPT-5.6 Sol Ultrafast hits 750 tokens/sec, answering 2,500 PhD-level questions in 11 hours—7x faster than rivals.

Deep Dive

OpenAI and Cerebras have launched **GPT-5.6 Sol Ultrafast**, a premium API tier designed to shatter the speed-intelligence tradeoff in AI inference. This service, powered by Cerebras’ Wafer-Scale Engine architecture, delivers up to **750 output tokens per second** without compromising model quality. The breakthrough is initially available to select OpenAI API customers, with broader access rolling out over time.

The innovation is validated by Cerebras’ internal benchmarks, including **Humanity’s Last Exam (HLE)**, a grueling 2,500-question benchmark requiring PhD-level expertise across fields like chemistry and economics. GPT-5.6 Sol Ultrafast answered all questions in **11 hours and 11 minutes**—a **7x speedup** compared to Claude Fable 5, which took over three days. On the GDP-Val benchmark for economically valuable knowledge work, Ultrafast achieved a **5.6x end-to-end speedup** with no quality loss, demonstrating its potential to transform high-stakes workflows like real-time cybersecurity incident response, legal document drafting, and financial modeling. Cerebras and OpenAI position Ultrafast as a critical tool for organizations needing AI to keep pace with human decision-making, enabling agents to operate in real time and reduce context-switching overhead for users.

Key Points
  • GPT-5.6 Sol Ultrafast delivers **750 tokens/sec** via OpenAI API, powered by Cerebras’ Wafer-Scale Engine—11x faster than Fable 5.
  • Answered **2,500 PhD-level questions in 11 hours** (7x faster than rivals) on Humanity’s Last Exam while maintaining accuracy.
  • Enables real-time agent workflows, cybersecurity threat detection, and legal/financial report generation with **no quality compromise**.
  • GDP-Val benchmark shows **5.6x speedup** for economically valuable tasks, proving productivity gains without tradeoffs.

Why It Matters

Ultrafast inference turns AI from a batch processor into a real-time collaborator, unlocking new productivity gains in high-stakes domains like cybersecurity and legal work.

📬 Get the top 10 AI stories daily