New UQ protocol for text-to-video models hits 98% accuracy, 1.00 AUROC
Wan2.1 14B's seed diversity is signal, not noise—new protocol decodes it
When a text-to-video model animates the same modern artwork with different random seeds, it produces visibly different films—one reading per seed. Because modern art is ambiguous by intent, this disagreement is signal, not noise. Yet existing uncertainty quantification (UQ) collapses the variations into a single dispersion scalar, unable to distinguish a compact interpretation from a dominant reading plus outlier, two competing modes, or diffuse instability. Tirtho Roy and colleagues from multiple institutions present the first systematic study of generative uncertainty for modern art animation, offering a reusable protocol that works with any text-to-video generator.
The protocol combines seven source-blind and six reference-aware estimators, a distributional profile covering robust spread, outlier influence, topology, multimodality, anisotropy, leave-one-seed influence, and reference coverage, plus a distribution model ablation (vMF, Kent, ACG, Student t, kernel, mixture). The team built a corpus of 250 modern artwork captions rendered by Wan2.1 14B under four seeds—1,000 videos total—across four encoding backbones, with artworks withheld from generation. The diagnostic succeeds dramatically: it classifies seed-set topology at balanced accuracy 0.98 (chance 0.25), isolates outlier configurations at AUROC 1.00 where a scalar reaches only 0.35, and splits high-uncertainty artworks into reference-covering (n=97) and reference-missing (n=56) diversity. Crucially, the protocol works reliably from just three seeds and across encoders, making it practical for real-world evaluation.
- Classifies seed-set topology at 98% balanced accuracy vs 25% chance across 1,000 Wan2.1 14B videos
- Isolates outlier configurations at AUROC 1.00, far exceeding scalar-based UQ (0.35)
- Protocol needs only 3 seeds and works across 4 encoders, splitting high-uncertainty artworks into coverage types
Why It Matters
Gives generative video teams a practical way to interpret seed diversity, not just measure it, improving evaluation of creative AI outputs.