FreeStory keeps characters consistent in AI visual stories without training
No more repeating full descriptions—FreeStory handles pronouns and free-form prompts
Visual storytelling with AI image generators often falters at keeping characters looking the same from one frame to the next. Existing training-free approaches solved this by forcing users to repeat a full character description in every prompt, which works for rigid scripts but breaks down in natural storytelling where a character is introduced once and later referred to as “he,” “she,” or “the detective.”
FreeStory, developed by Sibo Dong, Ismail Shaheen, and Sarah Adel Bargal, addresses this gap. It automatically associates references like pronouns with their original character descriptions, then uses dynamic masks, correspondence-aware feature matching, key-value injection, and query blending to preserve visual identity while allowing diverse poses and scenes. The team also created FreeStoryBench, a dedicated benchmark with single- and multi-character narratives written in free-form style. Experiments show FreeStory outperforms all existing training-free methods on structured benchmarks and achieves stronger consistency under truly free-form prompts, bringing AI storytelling closer to human-authored narratives.
- FreeStory handles free-form prompts with pronouns and type-based references instead of requiring repeated full descriptions
- Uses entity-grounded feature reuse combining dynamic masks, correspondence-aware matching, and attention blending
- Introduces FreeStoryBench benchmark with single and multi-character stories; achieves SOTA among training-free methods
Why It Matters
Enables more natural AI visual storytelling without extra training, making character consistency practical for real-world narratives.