Fixing 98% of Your Data Still Broke Every Decision
Data cleanup can look like a win while changing nothing that matters.
Here's what happened. A software researcher took two standard data-quality fixes — stamping every row with a label saying where it came from and how trustworthy it is, and forcing all data to enter through one controlled doorway — and pointed them at a real production system. The system had 194,620 rows of data and made two automated verification decisions. After the fixes, label coverage jumped from 36% to 98%, and 3,070 questionable entries were refused at the door. By every dashboard measure, it was a triumph.
Except it wasn't. Neither of the two decisions changed at all. Not for the better, not for the worse. Digging in, the researcher found the reason: the rows these fixes repaired and the rows the decisions actually read were completely different sets. Zero overlap. The single biggest fix moved 121,296 rows from "we can't tell where this came from" to properly labeled — and every one of those rows sat outside the window the decisions looked at. The author calls this an "empty intersection."
Why should you care? Because this exact pattern plays out in companies constantly. A team ships a data cleanup, the quality score goes green, everyone celebrates — and the actual report, fraud check, or AI recommendation still rests on the same shaky numbers it always did. The metric moved. The outcome didn't. Anyone who has sat through a quarterly review built on a repaired dashboard has felt a version of this.
The lesson isn't that data hygiene is pointless. Both fixes did exactly what they promised. The lesson is that they promised nothing about decisions. They described the whole population of data without saying which rows any decision would read. So before you pay for a cleanup, a compliance tool, or an AI data-integrity project, ask one unglamorous question: which specific records does the decision I care about actually touch? Fix those first. A perfect dataset nobody reads is just expensive tidying.
- A data fix raised labeling coverage from 36% to 98% — and changed zero outcomes it was supposed to protect
- 3,070 bad data entries were blocked at the door, yet the two decisions still had nothing trustworthy to work with
- The repaired rows and the 32 rows the decisions actually used never overlapped — fix what feeds the decision, not what pads the dashboard
Why It Matters
Before buying data cleanup or AI integrity tools, find out which records your key decisions actually read.