COSI-Lab dataset models social intent via multimodal conference recordings
32 academics, two mingling sessions, and a mission to decode hidden social intentions.
COSI-Lab is a new living-lab dataset that tackles a notoriously hard problem: how humans infer others' intentions during brief social encounters. Researchers recorded 32 academics during two 30-minute mingling sessions at an interdisciplinary workshop, where participants had genuine professional and social motivations—making the interactions ecologically valid rather than contrived. The team frames Apparent Intent Inference (AII) as a perspective-driven reasoning process: different observers may honestly interpret the same behavior differently, and that ambiguity should be modeled explicitly, not discarded as annotation noise.
The paper contributes four concrete assets: a novel annotation process that accounts for each perceiver's interpretative tendencies; quantitative and qualitative analyses of intent narratives (diversity, grounding, plausibility); benchmark tasks for AII and related contextual factors like social involvement; and privacy-preserving multimodal data including speech-quality audio for all participants, for downstream lexical and nonverbal behavior analysis. Uniquely, COSI-Lab couples participants' self-reported goals (spanning 30 minutes to 3 hours) with AII annotations at second-level granularity, enabling direct correlation between long-horizon intentions and moment-to-moment social perception. For AI researchers, this is a step beyond standard action recognition—it provides a substrate for building systems that can reason about why someone might be perceived as interested, disengaged, or persuasive in a dynamic social setting.
- Dataset captures 32 academics in two 30-minute weakly scripted mingling sessions with real professional and social consequences
- Introduces Apparent Intent Inference (AII) with an annotation process that models perceiver bias as explainable reasoning, not label noise
- Pairs self-reported participant goals (30-minute to 3-hour scope) with second-level AII annotations and privacy-preserving multimodal audio/video data
Why It Matters
Could help build context-aware AI assistants that understand unspoken social cues in professional networking and collaborative settings.