Research & Papers

Microsoft's VERDICT catches reasoning errors in multimodal AI, no training needed

Disagreement between verifiers signals bad reasoning steps, boosting accuracy by 5.95%

Deep Dive

Multimodal large language models (LLMs) often produce reasoning chains with subtle errors that lead to wrong answers. Existing verification methods either require expensive labeled supervision and fail to generalize, or naively aggregate scores from multiple sources, missing a critical signal: when verifiers disagree, that disagreement itself indicates an invalid reasoning step. Microsoft Research and collaborators from IIT Hyderabad introduce VERDICT, a training-free, domain-agnostic step-wise verifier that makes cross-modal disagreement explicit and actionable.

VERDICT formalizes verification as a coupled scoring problem among frozen, disparate verifiers, interpreted as a coordination game with a unique closed-form equilibrium. Agreement between verifiers signals valid steps, while disagreement reveals instability. This enables disagreement-aware filtering and stability-conscious ranking of reasoning steps. Evaluated across six benchmarks, VERDICT improves base model accuracy by up to +5.95%, performing competitively with heavily supervised, task-specific critics. Accepted at ECCV 2026, VERDICT shows that cross-modal agreement provides robust verification signals without task-specific adaptation, offering a practical path to more reliable multimodal reasoning.

Key Points
  • VERDICT is the first training-free verifier that explicitly models cross-modal disagreement for step-wise verification.
  • Uses a coordination-game formulation with a unique closed-form equilibrium, avoiding expensive labeled supervision.
  • Improves base multimodal models by up to +5.95% across six benchmarks, matching supervised critics.
  • Presented at ECCV 2026 by Microsoft Research and IIT Hyderabad (arXiv:2608.10665).

Why It Matters

Enables reliable verification of AI reasoning without training data, cutting costs and improving multimodal model trustworthiness.

📬 Get the top 10 AI stories daily