Agentic AI fails power grid tests: only 2 of 6 complexity levels solvable
Researchers tested agentic AI on 4 grid scales—and it couldn't handle complex planning
Power systems are buckling under the rapid growth of AI-driven data centers, which demand fast grid connections while transmission expansion lags. Agentic AI—autonomous systems that can plan and execute tasks—is proposed as a solution to automate repetitive connection processes. But according to a new arXiv study by Eve Tsybina, Samim Konjicija, and Slaven Peles, that promise remains largely unfulfilled. The researchers replicated current state-of-the-art agentic AI for power systems planning and tested it on a structured suite of nodal planning problems spanning six levels of task complexity and four grid scales.
The results are sobering: only the two lowest complexity levels were solvable, and only on some grid sizes. No agentic AI system could handle higher-complexity tasks or larger-scale networks, exposing a clear capability gap between current models and real-world operational needs. The authors argue that adopting stricter testing protocols and reproducible evaluation benchmarks is essential to measure genuine progress and operational readiness. Without these, they warn, agentic AI risks being deployed prematurely in critical infrastructure, with potentially costly consequences. The paper is short (5 pages, 1 figure, 2 tables) but delivers a pointed message: the industry needs standardized stress tests before trusting AI with the grid.
- State-of-the-art agentic AI solved only the 2 lowest complexity levels out of 6 in power systems planning tests.
- Evaluation covered 4 grid scales and found no agentic system handled large-scale networks reliably.
- Authors call for reproducible benchmarks and stricter testing protocols before operational deployment.
Why It Matters
Grid operators can't yet rely on agentic AI for complex decisions—benchmarks must come before deployment.