DeepSeek's V4-Flash outperforms premium models with 645% gain
DeepSeek's V4-Flash beats V4-Pro-Preview by 14.7% on agent tasks...
DeepSeek just redefined the agentic AI landscape with the public beta release of its DeepSeek-V4-Flash API (model ID: deepseek-v4-flash), a lightweight model that outperforms its premium V4-Pro-Preview counterpart across every agent benchmark despite identical architecture. The breakthrough comes from 'better post-training'βthe same model size but radically improved agentic capabilities. On Terminal Bench 2.1, V4-Flash scored 82.7 versus V4-Pro-Preview's 72.1, a 14.7% lead, while costing just $0.14 per million input tokens (cache-miss) compared to Anthropic's Opus-4.8, which scores 85.0 on the same benchmark but at significantly higher pricing.
The model's performance gains are staggering: a 645% improvement on DeepSWE (from 7.3 to 54.4) and dominant scores across NL2Repo (54.2), Cybergym (76.7), and Toolathlon Verified (70.3). DeepSeek also natively supports the OpenAI Responses API format and Codex integration, making it a seamless drop-in replacement for OpenAI models in agent pipelines. This positions V4-Flash as the most cost-effective frontier agent model, enabling teams to slash costs by up to 90% while maintaining high performance.
- V4-Flash outperforms V4-Pro-Preview by 14.7% on agent benchmarks using the same architecture, thanks to advanced post-training
- Priced at $0.14/M input tokens, it delivers 82.7 on Terminal Bench 2.1 versus Opus-4.8's 85.0 at a fraction of the cost
- Supports OpenAI Responses API natively and Codex integration, enabling drop-in replacements for OpenAI models
Why It Matters
AI teams can now deploy high-performance agent models at 90% lower cost, reshaping build-vs-buy decisions for automation workflows.