MultiRef-Compass benchmark tests AI audio-video generation with 14 metrics
First unified benchmark for multi-reference audio-video generation reveals major gaps in current models
Multi-reference-to-audio-video (MR2AV) generation is an emerging AI task where models create synchronized audio and video content conditioned on multiple reference inputs and textual instructions. Existing benchmarks focus on text-driven generation, single-reference subject preservation, or isolated audio-video alignment, leaving MR2AV largely unevaluated. To fill this gap, a team of researchers from multiple institutions developed MultiRef-Compass, a comprehensive benchmark designed specifically for MR2AV systems.
MultiRef-Compass comprises 350 samples built through a scalable asset-composition pipeline, covering scenarios like multi-view subject preservation, multi-entity binding, and human-object-scene composition. The evaluation protocol defines four dimensions: Basic Quality, Reference Consistency, Audio-Visual Consistency, and Instruction Following, measured via 14 sub-metrics. It integrates automatic metrics with a rejudging-enhanced MLLM-as-a-Judge framework for scalable and auditable evaluation. When tested on eight representative MR2AV systems, the benchmark revealed substantial room for improvement across all dimensions, positioning itself as a foundation for future research.
- 350 curated samples covering multi-view preservation, entity binding, and human-object-scene composition
- 14 sub-metrics across 4 evaluation dimensions: Basic Quality, Reference Consistency, Audio-Visual Consistency, Instruction Following
- Tested on 8 MR2AV systems; all showed significant gaps, underscoring the need for better models
Why It Matters
Drives progress in generating coherent audio-video content from multiple references, critical for advanced media production and AI-assisted creation.