Open Source

DeepSeek v4 Flash gets DS4 engine boost with DSpark MTP

DeepSeek v4 Flash + DS4 engine hits 30 tok/sec on M5 Max...

Deep Dive

DeepSeek's v4 Flash model has received a significant performance boost through antirez's DS4 (DwarfStar) inference engine, which now achieves 30+ tokens per second on an Apple M5 Max chip—double the speed of standard llama.cpp implementations. This comes after community developer returnity quickly released GGUF quantizations for the new checkpoint, though initial testing showed suboptimal speeds below 15 tok/sec in raw llama.cpp environments.

The DS4 engine's optimization includes persistent KV caching on SSD storage and compatibility with any coding harness via OpenAI API endpoints, reducing setup time to just 60 seconds. Additionally, returnity has released separate DSpark MTP head quantizations, enabling directional steering capabilities that can potentially de-censor the model. While antirez has since released his own 0731 quantizations, returnity's DSpark heads remain unique in their current form and are designed to work alongside them.

Key Points
  • DeepSeek v4 Flash + DS4 engine achieves 30+ tok/sec on M5 Max (2x faster than standard implementations)
  • New DSpark MTP head quantizations enable directional steering for uncensored outputs
  • DS4 supports SSD persistent KV caching and OpenAI API compatibility with 60-second setup

Why It Matters

Enables enterprise-grade AI agents with 2x faster inference and customizable uncensored outputs via DS4's unique features.

📬 Get the top 10 AI stories daily