Research & Papers

New Tool Makes Big Data Queries Up to 25x Faster

⚡The same analysis could finish in minutes instead of hours — and cost less.

Deep Dive

Meet Reffine, a compiler-based in-memory analytics engine from Anand Jayarajan and Gennady Pekhimenko that aims to deliver high performance across a broad range of data analytics workloads. Rather than hand-tuning each task, Reffine introduces a novel intermediate representation grounded in relational algebra and sparse iteration theory, giving data and computation a unified abstraction that enables workload-agnostic, end-to-end optimizations like operator fusion and automatic parallelization. A sparse compiler backend then translates that IR into hardware-efficient imperative code, aiming for high multi-core performance without domain-specific implementations. The results: on the TPC-H benchmark, Reffine outperforms the in-memory analytical database DuckDB by up to 24.9× and the state-of-the-art compilation-based database Umbra by up to 3.2×, and it achieves average speedups of 18.3× over Polars and 47.9× over NetworkX on streaming and graph analytics workloads, respectively. Source code is linked in the paper.

Key Points
  • Reffine is a free, research-built engine that makes in-memory data analysis dramatically faster by rewriting queries automatically — no hand-tuning per task.
  • It beat DuckDB by up to 25x on a standard industry benchmark, and averaged 18x faster than Polars and 48x faster than NetworkX on other workloads.
  • Faster analysis can mean lower cloud computing bills and quicker answers for reports, fraud checks, and recommendations — but this is a prototype, not a ready-made product.

Why It Matters

Faster data crunching could mean shorter waits and lower cloud bills for the apps you use daily.

📬 Get the top 10 AI stories daily