Research & Papers

CollabSkill: New benchmark ranks Claude Code #1 for human-AI collaboration

Real humans work alongside AI agents on real tasks—Claude Code beats Codex.

Deep Dive

A team of researchers (Yijia Shao, Zora Zhiruo Wang, et al.) has introduced CollabSkill, a framework designed to evaluate how well AI agents collaborate with human workers on real-world tasks. Instead of testing agents in isolation, CollabSkill pairs real workers with AI agents on tasks matched to their occupational background, capturing authentic usage patterns and the complexity of economically valuable work. The dataset includes over 1,500 prompts from 386 working sessions contributed by 93 human workers, and a Bayesian skill rating system disentangles the contributions of both humans and agents.

The results challenge conventional wisdom from fully autonomous benchmarks. While models like Codex lead in solo coding tasks, Claude Code ranks first in CollabSkill's human-agent evaluation. The study also reveals that hands-on collaboration experience significantly shifts workers' AI literacy, and practical experience is the biggest predictor of collaboration skill. This suggests that real-world human-AI teamwork requires different capabilities than standalone AI performance, and the community should invest in systematic evaluation of collaborative scenarios.

Key Points
  • 93 human workers with real occupational backgrounds contributed 1,500 prompts across 386 working sessions.
  • Claude Code ranks first in human-agent collaboration, while Codex leads in fully autonomous benchmarks.
  • Practical experience is the primary driver of collaboration skill, outweighing theoretical AI literacy.

Why It Matters

Shifts AI evaluation from isolated benchmarks to real-world human collaboration, revealing which agents truly augment workers.

📬 Get the top 10 AI stories daily