New AI Can Find Any Sound or Sight in Videos Without Training
This could make searching video as easy as searching text.
arXivLabs is a framework that lets collaborators develop and share new arXiv features directly on the arXiv website. Both individuals and organizations that work with arXivLabs have embraced and accepted arXiv's values of openness, community, excellence, and user data privacy.
arXiv is committed to these values and only works with partners who adhere to them. If you have an idea for a project that will add value for arXiv's community, you can learn more about arXivLabs.
- AI can now find specific sounds and objects in videos without prior training on those categories.
- It combines audio and visual cues using a novel 'complex-valued fusion' method for higher accuracy.
- Potential uses include video editing, surveillance, accessibility, and content moderation.
Why It Matters
This could make searching video as easy as searching text, saving time and enabling new applications.