Research & Papers

VIP-MINGLE dataset captures group language shifts between video and in-person settings

59 hours of paired recordings reveal behavioral differences across meeting modes

Deep Dive

A team of researchers from NYU, including Andrew Chang and David Poeppel, released VIP-MINGLE, a large-scale multimodal dataset designed to bridge the gap between videoconference and in-person group interactions. The dataset comprises 59 hours of recordings from 105 participants across 32 groups, each participating in paired sessions in both settings under the same lab conditions. It includes raw audio and video, psychometric surveys, processed multimodal features (e.g., speaker diarization, facial expressions, transcriptions), and time-resolved human annotations, making it a comprehensive resource for studying group language engagement.

Preliminary analysis of VIP-MINGLE reveals significant behavioral distribution shifts across multiple modalities—speech patterns, gaze, and turn-taking—between videoconference and in-person contexts. These shifts underscore the need for cross-setting corpora to avoid biased models. The dataset was accepted at Interspeech 2026 and is publicly available. It enables researchers to build robust AI models that generalize across meeting modes, potentially improving collaborative tools, virtual reality experiences, and social signal processing systems.

Key Points
  • 59 hours of recordings from 105 participants across 32 groups with paired in-person and videoconference sessions
  • Includes raw audio/video, psychometric data, diarized speech, facial expressions, and temporal annotations
  • Analysis shows significant behavioral shifts between settings, highlighting the need for cross-environment corpora

Why It Matters

Enables robust AI models for group dynamics across remote and in-person collaboration

📬 Get the top 10 AI stories daily