Research & Papers

Vision Transformers learn human-like visual grouping from natural images

ViTs innately grasp surroundedness and convexity from real-world scenes alone.

Deep Dive

Researchers found that Vision Transformers (ViTs) learn Gestalt-like figure-ground cues—surroundedness and convexity—directly from natural images. Testing 25 models (supervised and self-supervised) with linear probes, they observed robust encoding and zero-shot generalization to artificial stimuli. Symmetry was encoded only for uniformly colored regions, not textured ones. This positions ViTs as a compelling model system for studying the computational mechanisms of perceptual organization.

Key Points
  • ViTs robustly encode surroundedness and convexity from natural images, with zero-shot generalization to artificial stimuli.
  • Symmetry cue is learned only for uniformly colored regions, not textured ones, revealing a limitation in current models.
  • Study tested 25 ViTs across supervised and self-supervised objectives using linear probes on intermediate patch representations.

Why It Matters

Insights could improve AI vision for autonomous driving and medical imaging by mimicking human-like object-background separation.

📬 Get the top 10 AI stories daily