Models & Releases

DeepSeek V4 GA launches with 1M context, peak-valley pricing, challenges GPT-5.6 and Claude

DeepSeek V4 Pro at $0.87/1M tokens off-peak undercuts rivals by 10x or more.

Deep Dive

DeepSeek has officially launched DeepSeek V4 GA (general availability), rolling out the production version of its V4 family that had been in preview since April 2026. The lineup maintains two tiers: DeepSeek-V4-Pro with 1.6 trillion total parameters (49B active via MoE) and DeepSeek-V4-Flash with 284B total (13B active). Both models ship with a default 1 million token context window, dual thinking/non-thinking modes, and APIs compatible with OpenAI and Anthropic. The most notable change is the introduction of peak-valley pricing: API rates roughly double during Beijing peak hours (09:00–12:00 and 14:00–18:00 CST), while staying at baseline off-peak. For V4-Pro, output costs $0.87 per million tokens off-peak and $1.74 peak; for V4-Flash, $0.28 off-peak and $0.56 peak. Cache-hit input on Flash can be as low as $0.0028/1M tokens. This makes DeepSeek dramatically cheaper than US flagships even at peak pricing.

Compared to rivals, DeepSeek V4 Pro sits near Claude Opus 4.8/Fable 5 territory in quality, with strong coding, agent capabilities, and SVG/3D generation, though Claude still wins on long-horizon software engineering and retrieval-heavy tasks. Against GPT-5.5/5.6, DeepSeek trails slightly on agent benchmarks but competes hard on coding and math, and undercuts GPT API spend by a wide margin. Google Gemini 3.1 Pro still leads for deep multimodal and Google-ecosystem workflows, but DeepSeek’s open weights and 1M context make it more attractive for self-hosting, privacy-sensitive stacks, and cost-controlled RAG. For high-volume extraction, summarization, and first-draft code, V4 Flash or Pro is often the rational default. Legacy model IDs (deepseek-chat, deepseek-reasoner) retire after July 24, 2026.

Key Points
  • DeepSeek V4 GA includes two models: V4-Pro (1.6T params, 49B active) and V4-Flash (284B, 13B active), both with 1M context.
  • New peak-valley pricing: V4-Pro output at $0.87/1M tokens off-peak, $1.74 peak; Flash at $0.28/$0.56; cache hits as low as $0.0028.
  • Comparable to Claude Fable 5 and GPT-5.6 on many tasks, but 10-100x cheaper; strong for coding, math, and self-hosted RAG.

Why It Matters

DeepSeek V4 GA resets AI economics: production-grade intelligence at a fraction of US rival costs, especially for volume and self-hosted workflows.

📬 Get the top 10 AI stories daily