CARE v1.0 dataset: 144 hours of multimodal interviews across 12 medical conditions
New 144-hour multimodal corpus to advance AI-driven disease detection and monitoring.
CARE v1.0, introduced by Gimeno-Gómez et al., is a curated multimodal English dataset comprising approximately 144 hours of short video interviews collected from 612 individuals across 12 distinct medical conditions—including neurological, psychiatric, and respiratory disorders—plus a healthy control cohort. Each video is annotated with a comprehensive set of clinically relevant multimodal descriptors covering speech, non-verbal communication (e.g., facial expressions, gestures), and rich metadata such as medication use, comorbidities, life impacts, and expressed emotions. This addresses a critical gap: existing publicly accessible datasets for automated analysis of medical speech are often small, single-condition, and primarily audio-focused, with insufficient documentation of confounding variables like education or mood state.
The corpus supports a wide range of applications: automatic disease and symptom detection, multimodal modeling of speech and non-verbal behavior under emotionally charged contexts, and longitudinal studies of disease progression and coping mechanisms. The dataset’s heterogeneity and breadth enable more robust and generalizable AI models for healthcare. Researchers are calling for community collaboration to further extend and validate use cases. The paper is currently under review at npj Scientific Data and the dataset is expected to be publicly released to accelerate research in computational phenotyping and assistive technologies.
- 144 hours of video interviews from 612 individuals across 12 medical conditions plus a control group.
- Rich multimodal descriptors including speech, non-verbal behavior, medication, life impacts, and emotions.
- Addresses lack of large-scale, multi-condition, publicly available datasets for AI-driven health monitoring.
Why It Matters
CARE v1.0 unlocks AI-driven health monitoring across 12 conditions, enabling robust multimodal research and real-world clinical applications.