Research & Papers

Flink vs Kafka Streams: A New Study Shows Which One Survives a Crash

⚡The engine behind instant payments and fraud alerts can quietly stall for 23 seconds.

Deep Dive

Researchers ran a controlled comparison of exactly-once Kafka pipelines built with Apache Flink and Kafka Streams — two engines that manage state very differently: Flink checkpoints to external storage, while Kafka Streams restores local state by replaying broker changelogs. Across 50 correctness trials (five for each engine-workload combination), both engines produced expected outputs, with an exact 95% CI of 0.929–1.000. Speed told a different story. In 30-minute trials at a fixed 100 events/sec, the measured ingestion-to-output interval was 4–6 ms for stateless workloads but 4.3–6.2 s for windowed ones, and the median p99 was 5.68–9.92 s for windowed workloads. A sensitivity study that raised the durability interval from 1,000 to 10,000 ms shifted latency distributions: Kafka Streams' W3 median p99 T2–T1 climbed from 6,085 to 23,155 ms, while Flink's T3–T2 rose from 732 to 8,702 ms. Under injected failures, Kafka Streams completed 0 of 5 stateful JVM-kill trials and 1 of 5 in each stateful

Key Points
  • Both tools were 100% accurate in 50 trials, but accuracy alone hid big differences in speed and reliability.
  • Changing one setting — how often the system saves its progress — pushed worst-case delays from about 6 seconds to 23 seconds.
  • When researchers forced crashes, Flink recovered every time while Kafka Streams failed all five process-kill tests.

Why It Matters

If your bank picks the wrong real-time engine, fraud alerts can lag 20 seconds or vanish in a crash.

📬 Get the top 10 AI stories daily