Startups & Funding

OpenAI's Ultrafast mode runs GPT-5.6 Sol at 14x speed

Ultrafast hits 750 output tokens per second — 14x faster than standard ChatGPT.

Deep Dive

OpenAI has introduced Ultrafast, a new processing mode designed specifically for its most powerful model, GPT-5.6 Sol. The mode reportedly runs at 14x the speed of standard inference, generating up to 750 output tokens per second. In its Thursday announcement, OpenAI framed this as a breakthrough in efficiency: “Until now, getting real-time speed typically meant choosing a smaller or more specialized model. Ultrafast points to progress in a new direction: more useful work per second.” The feature is powered by a partnership with AI chip maker Cerebras, whose specialized hardware enables the dramatic throughput boost.

The preview is currently available to a small group of customers, with OpenAI stating it will expand access as capacity grows. The company suggests Ultrafast is ideal for heavy corporate workloads that demand low latency and high throughput, including incident response, customer service and support, real-time financial market analysis, and e-commerce. While Anthropic's Claude offers a "fast mode," OpenAI's implementation delivers significantly higher raw speed. Ultrafast marks a strategic shift toward using advanced models at scale without sacrificing responsiveness, putting large language models on par with the speed of smaller, task-specific alternatives.

Key Points
  • OpenAI's Ultrafast mode accelerates GPT-5.6 Sol to 14x standard speed, yielding up to 750 output tokens per second.
  • The mode is powered by an OpenAI partnership with chipmaker Cerebras, using specialized hardware for high-throughput inference.
  • Currently in preview for select customers, with target workflows including incident response, customer support, and financial market analysis.

Why It Matters

Ultrafast lets enterprises deploy top-tier LLMs in real-time operations, closing the latency gap with smaller models.

📬 Get the top 10 AI stories daily