CaM-Wolf: First multimodal AI agent masters Werewolf with video input
CaM-Wolf reads player faces and uses causal reasoning to dominate social deduction games.
Researchers have unveiled CaM-Wolf, a groundbreaking AI agent designed for social deduction games (SDGs) like Werewolf, accepted at ACMMM 2026. While existing LLM-powered agents operate purely on text, CaM-Wolf integrates multimodal perception—processing live video feeds of other players to capture facial expressions, gestures, and other visual cues. This allows the agent to go beyond reading chat logs and actually "see" the social dynamics at the table.
At the core of CaM-Wolf is a causal-aware Reasoner trained via reinforcement learning. This module establishes logical chains between observed player behaviors (e.g., nervous glances, confident posture) and hidden roles (werewolf, villager, seer). The agent then presents itself through an animated avatar, enabling rich, non-verbal communication during gameplay. Experiments and user studies demonstrate that CaM-Wolf achieves superior performance in winning games and significantly improves the quality of human-AI social interaction, making it a major leap toward AI that can participate in nuanced, real-world social scenarios.
- CaM-Wolf is the first multimodal agent for social deduction games, processing video inputs from players in real time.
- Uses a causal-aware Reasoner trained via reinforcement learning to link observable behaviors to hidden roles.
- Outperforms text-only agents in gameplay and improves human-AI interaction quality, as shown in experiments and user studies.
Why It Matters
Brings AI closer to human-like social reasoning by integrating visual cues and causal logic into game-playing agents.