Sirotkina's study: AI saliency models beat by simple center, biased by demographics
11.4M gaze points show trained networks lose to an untrained center marker.
Computer vision saliency models promise to predict where people look, powering a billion-dollar attention-prediction industry. But according to a new arXiv paper by Elena Sirotkina, those models barely beat a naive baseline and carry hidden demographic biases. She tested leading saliency networks against 11.4 million real webcam gaze points collected from 3,023 US adults — recruited to national quotas — while they viewed circulating news photographs. The result: an untrained central marker (simply always predicting the image center) outperformed every trained network. The models' added content falls exactly where these audiences never look.
The study also exposes systematic bias in whatever accuracy remains. The models' predictions align better with younger, White, and politically moderate viewers, while performing worse for older, Black, and ideologically extreme participants. To fix this, Sirotkina introduces a new evaluative framework built on what a group's own gaze reveals about whether a model can learn that group at all. She applies it across demographic axes and argues that systems deciding what people see can learn to see everyone — but only if the industry adopts rigorous, gaze-based benchmarks. This paper supplies the standard for judging such claims.
- An untrained center marker beat every trained saliency network tested against 11.4M real gaze points.
- Saliency model accuracy is biased: it favors younger, White, and moderate viewers over older, Black, and ideologically extreme ones.
- Sirotkina proposes a new gaze-based evaluation framework to measure whether a model can learn any demographic group's attention patterns.
Why It Matters
If billion-dollar attention prediction fails without bias, advertisers and platforms need new standards or risk systematically misseeing audiences.