Research & Papers

Counterfactual fairness doesn't guarantee group fairness in image classifiers

New study reveals hidden biases in AI image classifiers, challenging prior assumptions.

Deep Dive

A new paper from researchers at Korea University and NAVER AI Lab challenges conventional wisdom about algorithmic fairness in computer vision. The team—Sangwon Jung, Sumin Yu, Sanghyuk Chun, and Taesup Moon—demonstrates that counterfactual fairness (CF) does not automatically lead to group fairness (GF) in image classifiers, a stark departure from results previously observed in tabular datasets.

To conduct their study, the authors constructed two new benchmark datasets (oursceleb and ourslfw) by applying high-quality image editing techniques and human annotation to create counterfactual samples—e.g., altering secondary sex characteristics while preserving identity. This allowed simultaneous evaluation of CF and GF. Their theoretical analysis reveals that a latent attribute G (such as hair length, which is correlated with but not caused by sex) introduces confounding, causing CF–achieving models to still exhibit group disparities. They propose Counterfactual Knowledge Distillation (CKD) as a simple baseline that reduces the model's reliance on G, enabling both fairness notions to align.

Key Points
  • Contrary to tabular data, counterfactual fairness does not imply group fairness in image classification.
  • Researchers created new datasets (oursceleb, ourslfw) using AI image editing to evaluate both fairness types simultaneously.
  • Proposed Counterfactual Knowledge Distillation (CKD) mitigates latent attribute confounding, achieving both fairness goals.

Why It Matters

This work reveals a hidden source of bias in vision AI, offering a practical fix for fairer deployment.

📬 Get the top 10 AI stories daily