KDD 2026 paper proposes statistical auditing for deployed AI systems
Treating fairness and safety as risk constraints under uncertainty in the wild.
Traditional AI evaluation relies on static benchmarks and sandboxed testing, which fail to capture real-world behavior shaped by shifting data distributions, user interactions, and infrastructure changes. The authors argue this limited view leaves deployed systems unmonitored, risking fairness and safety failures that cannot be caught by pre-deployment validation alone.
Their solution redefines auditing as continuous monitoring of risk-controlled constraints under statistical uncertainty. This approach treats properties like fairness and safety as constraints that must hold with high probability as the system evolves. They call for new monitoring algorithms that account for uncertainty, socio-technical criteria that go beyond narrow metrics, and infrastructure enabling long-term oversight. The work was accepted to the KDD 2026 Blue Sky Ideas Track.
- Traditional ML evaluation misses real-world behavior due to dynamic environments and user interactions.
- Proposes auditing as a statistical problem of monitoring constraint violations under uncertainty.
- Accepted to KDD 2026 Blue Sky Track, emphasizing socio-technical specs and oversight infrastructure.
Why It Matters
As AI systems are deployed widely, this framework provides a path to ensure ongoing fairness and safety.