Research & Papers

New research reveals calibration flaws in human-AI teams

Combining human and AI predictions breaks calibration — delegation shifts the burden.

Deep Dive

A new paper from Eric Nalisnick, Chi Zhang, Sophia Qian, and Yixin Wang tackles a fundamental question in human-AI collaboration: when a human and an AI model work together, how do their individual calibration properties affect the team's reliability? The study, submitted to arXiv on June 9, 2026, examines two common teaming frameworks: combination (averaging or mixing predictions) and delegation (handing off decisions to either the human or the model). The authors assume both the human and the AI are calibrated with respect to some partitioning of the feature space, then trace how those calibration guarantees propagate.

Their theoretical and empirical results reveal a stark trade-off. Combination methods never preserve the human's original calibration, meaning the team's confidence scores become unreliable even if each member was well-calibrated individually. Delegation frameworks do preserve calibration for the downstream predictor (the one who actually makes the decision), but they shift the burden onto the rejector — the meta-model that decides who should predict. That rejector must be calibrated at a much finer granularity to determine where each team member excels. Worse, when the human has private information the AI cannot observe, the rejector's task becomes impossible, creating a fundamental limit on delegation-based human-AI teams.

Key Points
  • Combining human & AI predictions breaks the human's calibration, reducing team reliability.
  • Delegation preserves calibration of the active predictor but requires a highly calibrated rejector.
  • When humans use unobserved information, the rejector's calibration demand becomes unattainable.

Why It Matters

Highlights critical design choices for building trustworthy human-AI systems in high-stakes domains.

📬 Get the top 10 AI stories daily