Developer Tools

LegacyWorld benchmark finds GUI agents leave dangerous side effects in 28 legacy workflows

Six computer-use agents tested on Windows—failed runs leave persistent invalid changes

Deep Dive

A new study from researchers at TU Munich introduces LegacyWorld, an atomicity-aware benchmark for evaluating GUI agents on legacy enterprise workflows. Legacy systems often lack programmable interfaces, forcing manual GUI interaction—and modern multimodal LLM agents (computer-use agents) promise to automate these tasks. But the paper warns that successful demos aren't enough: when an agent fails mid-run, it may leave persistent invalid changes in business or healthcare records.

To address this, the team built 28 Windows GUI workflows, each with an initial state, goal state, and task-specific validator. They compared expert-crafted prompts against prompts generated from screen recordings of expert golden-path executions, testing six hosted computer-use agents. Results show three distinct operational profiles: useful completion, safe failure, and non-atomic side effects. The authors argue that workflow capture, state validators, and atomicity-aware acceptance tests must become first-class requirements for AI-based legacy workflow automation, not afterthoughts.

Key Points
  • LegacyWorld benchmark covers 28 stateful Windows GUI workflows with task-specific validators
  • Six hosted computer-use agents were evaluated comparing expert prompts vs. screen-recording-generated prompts
  • Study reveals failed agent runs often leave non-atomic side effects, distinct from safe failures
  • Authors: Thilo Reintjes, Sivajeet Chand, Derui Zhu, Sushant Kumar Pandey, Alexander Pretschner (TU Munich)

Why It Matters

As AI agents automate legacy systems, atomicity—ensuring failed runs cause no persistent harm—becomes critical for real enterprise and healthcare deployments.

📬 Get the top 10 AI stories daily