Models & Releases

OpenAI's GPT-4o mini beats Gemini Flash and Claude Haiku on benchmarks at 60% lower cost

128K context, 82% MMLU, and 60% cheaper than GPT-3.5 Turbo — all in a compact model.

Deep Dive

OpenAI has released GPT-4o mini, a smaller and far more affordable model designed to compete directly with Google's Gemini 1.5 Flash and Anthropic's Claude 3 Haiku. Despite its size, GPT-4o mini outperforms GPT-4 on many tasks, scoring 82% on the MMLU benchmark — significantly ahead of Gemini Flash (77.9%) and Claude Haiku (73.8%). The model features a 128,000-token context window and supports up to 16,000 output tokens per request, with knowledge updated through October 2023.

Pricing is aggressive: $0.15 per million input tokens and $0.60 per million output tokens, making it 60% cheaper than GPT-3.5 Turbo. For comparison, Claude 3 Haiku costs $0.25 per million input tokens and $1.25 per million output tokens. Initially, GPT-4o mini supports text and vision in the API via Assistants, Chat Completions, and Batch APIs. Full multimodal support (text, image, video, audio) will arrive later. For end users, it's available immediately for Free, Plus, and Team subscribers, with Enterprise access rolling out next week. Fine-tuning is planned soon, giving developers even greater flexibility.

Key Points
  • GPT-4o mini scores 82% on MMLU, beating Gemini Flash (77.9%) and Claude Haiku (73.8%)
  • Priced at $0.15/M input tokens and $0.60/M output tokens — 60% cheaper than GPT-3.5 Turbo
  • Supports 128K context window, 16K output tokens, with text and vision API now; full multimodal later

Why It Matters

Developers get a high-performance, low-cost AI model that makes advanced language capabilities accessible for mass-market apps.

📬 Get the top 10 AI stories daily