Audio & Speech

AV-Flamingo: Open model understands long videos with audio and visual reasoning

Trained on 7M instances, it outperforms much larger closed models on complex video tasks.

Deep Dive

Audio-Visual Flamingo (AV-Flamingo) is a fully open, state-of-the-art audio-visual large language model designed for joint understanding and reasoning over audio, images, and long-form videos. Unlike prior models that focus on short clips, AV-Flamingo handles complex, real-world videos lasting minutes or more. The model introduces three key innovations: Audio-Visual-Skills, a dataset of ~7M caption and question-answer training instances emphasizing temporal, compositional, and cross-modal reasoning; a three-stage curriculum that progressively trains from short-range perception to long-horizon multi-event reasoning; and Temporal Audio-Visual Interleaved Chain-of-Thought, which explicitly grounds intermediate reasoning steps to timestamps, improving temporal alignment and interpretability.

Extensive experiments across 15+ audio-visual, omni-modal, audio, and vision benchmarks show that AV-Flamingo outperforms similarly sized open models by clear margins and remains highly competitive with—and in some cases surpasses—much larger open-weight and closed models, particularly on long and complex real-world audio-visual understanding and reasoning tasks. Beyond benchmarks, AV-Flamingo demonstrates strong real-world utility and transfers well to unseen tasks, highlighting its robustness and generalization ability. The model is fully open, with training data, code, and weights to be released, enabling researchers and developers to build on top of it.

Key Points
  • Trained on 7M caption and QA instances from real-world long videos, emphasizing temporal and cross-modal reasoning.
  • Uses a novel three-stage curriculum to progressively train from short-range perception to long-horizon multi-event reasoning.
  • Introduces Temporal Audio-Visual Interleaved Chain-of-Thought, grounding reasoning steps to specific timestamps for interpretability.

Why It Matters

Enables AI to deeply understand long, complex videos for surveillance, content moderation, and accessibility applications.

📬 Get the top 10 AI stories daily