Research & Papers

Researchers propose GazeAnywhere model for flexible gaze tracking

GazeAnywhere uses text prompts to estimate gaze targets anywhere in images

Deep Dive

Researchers from Georgia Institute of Technology and Meta AI have proposed GazeAnywhere, a transformer-based model that redefines gaze target estimation by introducing the Promptable Gaze Target Estimation (PGE) paradigm. This new approach allows users to specify gaze targets using natural language prompts—such as “the boy in the red shirt” or “person at coordinates [0.52, 0.48]”—eliminating the need for rigid, multi-stage pipelines that rely on head bounding boxes or pose estimation.

The team also introduced Gaze-Co, a 120,000-image dataset with prompt annotations, generated using a scalable data engine. GazeAnywhere integrates subject localization, in/out-of-frame detection, and gaze heatmap prediction into a single end-to-end model. In evaluations, it outperforms prior methods on multiple PGE benchmarks and maintains strong performance on out-of-domain clinical datasets. The model and benchmark are open-sourced and available on GitHub.

Key Points
  • GazeAnywhere is the first model designed for PGE, enabling flexible gaze target estimation via text or visual prompts
  • Trained on 120K images in the Gaze-Co dataset, it uses a transformer to fuse features and predict gaze heatmaps end-to-end
  • Achieves state-of-the-art results on multiple benchmarks and supports open-source release via GitHub

Why It Matters

Enables scalable, flexible gaze analysis in real-world settings without brittle multi-stage pipelines

📬 Get the top 10 AI stories daily