AI Video Models Just Got 30x Faster And Smarter
Your future AI video assistant could respond 30x faster, saving you minutes every day
A new arXiv paper introduces OraRL, a reinforcement learning approach that treats annotations as oracle rollouts to make post-training for video MLLMs more efficient and scalable. It tackles a failure called "advantage inversion" with a decoupled advantage estimator and sign-balanced pruning. OraRL needs only 2.2x the step time of SFT—less than half the 4.9x required by GRPO with chain-of-thought—and scales from 0.8B to 9B, surpassing GRPO up to 100k prompts. Without chain-of-thought, Video-ORA-9B decodes in 130 ms instead of 4,780 ms. It also lifts temporal mIoU from 62.5 to 66.0, tracking AO from 73.0 to 78.2, segmentation from 64.3 to 70.4, and the spatial-intelligence macro average from 51.0 to 56.1, scoring 73.1 on VSI-Bench versus 55.0 for GPT-5 and 55.1 for Gemini-3-Pro.
- AI video tools just became 30x faster thanks to a new training method called OraRL
- The tech improves accuracy in tracking objects and understanding video content
- Faster AI video analysis could soon appear in apps you use every day
Why It Matters
Expect near-instant video editing, smarter security cameras, and quicker self-driving car responses in the next few years.