DeepSeek-V4-Flash jumps 25 points on benchmarks
DeepSeek's V4-Flash jumps 25+ points on key benchmarks, rivaling GPT-5.6 Terra
DeepSeek has pushed its V4-Flash model into new territory with a significant performance leap that’s turning heads in AI circles. The update—dubbed a "HUGE improvement" by contributors—catapults the model’s Terminal Bench score from 56.9 to 82.7 (+25.8) and Toolathlon from 51.8 to 70.3 (+18.5). These gains make it competitive with OpenAI’s GPT-5.6 Terra on some benchmarks, though with stark differences in others (e.g., Flash leads in Terminal Bench and Toolathlon but lags in DeepSWE and Agents' Last Exam). The model’s 284B-A13B parameter architecture remains lightweight enough to run locally on under $10K hardware or even a single M5 Max, a rarity for models of this caliber.
The update reflects DeepSeek’s strategy of rapid iteration and post-training refinement, leveraging real-world developer data from providers like OpenRouter to improve capabilities. While cheaper than rivals, V4-Flash lacks vision capabilities—a trade-off noted by users for agentic tasks. The community’s excitement stems from its affordability, local inference potential, and the model’s ability to "get good enough" for broader use cases. With V4-Pro’s general availability on the horizon, DeepSeek is positioning itself as a disruptor in the cost-performance frontier of AI models.
- DeepSeek-V4-Flash’s Terminal Bench score jumped 25.8 points (56.9→82.7) and Toolathlon by 18.5 points (51.8→70.3) in the latest update.
- The 284B-A13B model competes with OpenAI’s GPT-5.6 Terra on some metrics while costing significantly less and running locally on under $10K hardware.
- Lacks vision capabilities but excels in terminal-based and tool-use benchmarks, making it ideal for cost-sensitive, agentic workflows.
Why It Matters
DeepSeek-V4-Flash delivers near-SOTA performance at a fraction of the cost, democratizing high-performance AI for developers and enterprises alike.