Research & Papers

New tumbling-E benchmark measures temporal reliability for human-AI comparison

1,154 trials reveal human perception timing with 1.5-second reaction time baseline.

Deep Dive

A new preprint from researchers Avneet Sandhu and Bin Hu introduces a dynamic computerized tumbling-E test designed to measure not just static visual acuity but the temporal reliability of human sequential perceptual decisions. Unlike traditional chart-based tests that yield a single acuity score, this adaptive staircase task records reaction times, timeout rates, and trial-level adaptation patterns. Over 1,154 valid trials from 21 participants across 77 sessions, the study captured a mean reaction time of 1,546 milliseconds, a timeout rate of 6.6%, and a smooth adaptation from 29.42 arcminutes at the start to 5.04 arcminutes by trial 19—converging near 20/20 level. These results demonstrate fast, reliable human timing and provide a richly annotated dataset of how perception unfolds over multiple decisions.

The key contribution is the introduction of the Temporal Hallucination Index (THI), which quantifies how static accuracy can mask delays, drift, persistence, and unstable convergence in decision-making. By converting the classic tumbling-E task into a temporally resolved benchmark, the authors create a human-only baseline for future comparison with artificial agents. This matters for AI safety and alignment: as AI systems increasingly make sequential perceptual decisions (e.g., in autonomous driving or medical imaging), understanding human timing patterns helps identify where models may appear accurate but deviate in temporal reliability. The dataset is publicly available for researchers working on human-AI comparison in perception tasks.

Key Points
  • Dataset includes 1,154 trials from 21 human identifiers across 77 sessions with 1,078 non-timeout responses and 76 timeouts (6.6% timeout rate).
  • Mean non-timeout reaction time was 1,546 ms with IQR 1,306–1,713 ms; only 3 responses exceeded 3,000 ms.
  • Stimulus size adapted from 29.42 to 5.04 arcminutes by trial 19, with 89.2% transitions to smaller optotypes, converging near 20/20 vision.

Why It Matters

Establishes a temporal reliability baseline for human perception, enabling fairer comparison with AI agents in sequential decision tasks.

📬 Get the top 10 AI stories daily