Open Source

DeepSeek V4 Flash 0731 hits 200 tps prompt, 11 tps gen on 4x RTX 5060 Ti

Reddit user reports 200 tps prompt processing on DeepSeek V4 Flash 0731 with 4x 5060 Ti GPUs

Deep Dive

DeepSeek V4 Flash 0731 is being tested by local AI enthusiasts, with one Reddit user posting real-world performance numbers from a 4x RTX 5060 Ti 16GB setup. Using llama.cpp with Unsloth's Q8 lossless quantization, they measured roughly 200 tokens per second for prompt processing and around 11 tokens per second for token generation. The system also features DDR4 3200 memory running in 4-channel mode, a 128K context window, and batch sizes (-ub/-b) set to 4096, providing a clear picture of what this configuration can handle.

These numbers are significant because DeepSeek V4 Flash 0731 is a large, competitive model that typically requires serious datacenter hardware. Running it effectively on four consumer GPUs opens the door for privacy-conscious teams, offline deployments, and cost-sensitive AI projects. The Q8 quantization retains near full precision, meaning users get strong output quality without sacrificing too much speed. While 11 tps is slower than cloud APIs, it's usable for many automation and research tasks, and prompt processing at 200 tps keeps interactive workloads responsive.

Key Points
  • ~200 tps prompt processing and ~11 tps generation on 4x RTX 5060 Ti 16GB via llama.cpp
  • Uses Q8 lossless quantization from Unsloth with 128K context window and 4096 batch size
  • DDR4 3200 RAM in 4-channel mode enables viable local inference of a frontier-style model

Why It Matters

Shows DeepSeek V4 Flash 0731 can run locally on consumer GPUs, enabling private, cost-effective inference without cloud dependencies.

📬 Get the top 10 AI stories daily