Research & Papers

ADEPT agent uses entropy to overcome ambiguous video retrieval queries

A training-free AI that asks clarifying questions or refines your query to find the right video.

Deep Dive

The core challenge in video retrieval from massive datasets is the 'Intent-Query Gap': users often cannot express their intent in a single text query. Traditional single-round methods lack feedback mechanisms, leading to poor results for complex or ambiguous searches. This paper proposes ADEPT, an entropy-driven agent that solves this by maintaining an interactive dialogue with the user.

ADEPT uses an entropy-based decision engine to choose between two strategies: ASK (asking the user clarifying questions to reduce uncertainty) and REFINE (automatically adjusting the query based on current results). The agent requires no training—it works out-of-the-box. Experiments on two challenging datasets show ADEPT outperforms all non-interactive, heuristic, and even Video-LLM baselines, setting a new performance benchmark. Published at IEEE ICASSP 2026.

Key Points
  • ADEPT uses an entropy-driven decision engine to dynamically choose between ASK and REFINE strategies.
  • The agent is training-free and requires no fine-tuning, working immediately on any dataset.
  • Outperforms all baselines on two challenging video retrieval benchmarks, including Video-LLMs.

Why It Matters

Makes massive video search practical for vague queries, eliminating the need for exact wording.

📬 Get the top 10 AI stories daily