Sidekick: New AI system boosts multitasking with multimodal feedback
30 participants saw 40% better multitasking with Sidekick's ambient cues and summaries.
Sidekick addresses a critical bottleneck in human-AI collaboration: the cognitive load of monitoring AI agents that automate GUI tasks. Current CUAs rely on text logs, forcing users to constantly shift attention to track progress. The research team—Ruei-Che Chang, Wenqian Xu, Dingzeyu Li, Bryan Wang, and Anhong Guo—designed Sidekick with three interaction modes: ambient cues (e.g., color changes, gentle sounds) for background status, multimodal summaries (visual + text) when returning to the agent, and real-time verbalization of reasoning when the agent operates in the foreground.
A controlled study with 30 participants showed Sidekick's multimodal approach dramatically improved multitasking efficiency. Users reported higher situation awareness and could more quickly detect errors and trace actions. The paper, published on arXiv (arXiv:2607.17527), positions Sidekick as a template for long-horizon human-agent collaboration, with implications for productivity tools, accessibility, and any domain where users need to delegate complex GUI tasks without constant supervision.
- Sidekick uses ambient cues, multimodal summaries, and verbalized reasoning to reduce cognitive load when monitoring AI agents.
- 30 participants in the study showed significant improvements in multitasking performance and error traceability over text-only baselines.
- The prototype is designed for Computer Use Agents (CUAs) that automate multi-step GUI tasks autonomously.
Why It Matters
Sidekick's design makes AI agents more practical for real-world multitasking, reducing the need for constant oversight.