Scientists Tested 29 Top AIs on Reading Brain Waves — Most Failed
Brain-reading AI could speed up sleep and epilepsy diagnoses — but it isn't ready yet.
Researchers built EEGAgentBench, a unified benchmark for systematically evaluating LLM agents on short- and long-horizon EEG analysis — a field where existing agentic evaluations were fragmented, covering limited tasks over narrow time horizons with inconsistent protocols. The benchmark spans six representative EEG applications, from knowledge question answering to sleep staging, with signal durations from 2 seconds to nearly 23 hours and prediction targets ranging from class labels to event intervals and epoch-level sequences. It provides 10 deterministic EEG analysis tools that expose only task-relevant signal measurements, so agents must select tools autonomously, accumulate evidence iteratively, and construct multi-step workflows. Benchmarking 29 frontier LLMs from 15 model families showed EEGAgentBench distinguishes agent capabilities beyond model scale and inference cost, while revealing substantial limitations of current LLM agents in long-horizon EEG analysis, particularly in sustained evidence accumulation and multi-step reasoning.
- EEG (brain-wave recordings) helps diagnose sleep problems and epilepsy, but someone has to read hours of squiggly lines by hand.
- Researchers tested 29 top AI models on six brain-reading tasks, from 2-second clips to nearly 23-hour overnight recordings.
- Most AI handled short clips, but nearly all failed on long recordings — and bigger, pricier models weren't any better.
Why It Matters
One day AI could read overnight brain recordings for doctors, speeding up sleep and epilepsy diagnoses — just not yet.