Research & Papers

LCG framework delivers consistent characters across 20-image sequences

Sparse attention mechanism keeps characters consistent from frame one to twenty.

Deep Dive

A team of researchers (Zihao Wang, Yijia Xu, Haoze Zheng, Xuran Ma, Haokun Gui, Harry Yang) introduced LCG (Long-Context Generation), a framework designed to solve a persistent problem in image generation: keeping characters and scenes consistent across a sequence of images. While single-image generation has advanced rapidly, models often fail to maintain visual coherence when asked to produce a series of related images for comics, storyboards, or visual narratives. LCG tackles this with two key innovations: Sparse Relational Attention (SRA) and Routing Consistency Constraint (RCC). SRA selectively attends to the most critical semantic and layout features across a long context, keeping computation tractable even as the number of images grows. RCC uses identity-aware masks to align structural patterns across generation branches, effectively preventing character appearance drift in complex multi-character scenes.

To train and benchmark their approach, the team built a large-scale synthetic dataset called LCCD (Long-Context Consistency Dataset). It contains 600,000 training sequences and a separate 1,000-test set, each sequence ranging from 6 to 20 images centered on characters in varied situations. In experiments, LCG outperformed baseline methods in both prompt alignment and character consistency — even for scenes with multiple characters. The implications are significant for professionals working in narrative visual generation: LCG offers a scalable, computationally efficient way to generate coherent visual stories, automatically ensuring that a character seen in the first frame still looks the same by the twentieth.

Key Points
  • LCG uses Sparse Relational Attention (SRA) to selectively attend to core features across long visual contexts, making multi-image generation computationally tractable.
  • The Routing Consistency Constraint (RCC) enforces identity-aware mask alignment to prevent appearance drift, even in complex multi-character scenes.
  • The LCCD dataset provides 600K training sequences of 6–20 images each, enabling rigorous training and evaluation for long-context consistency.

Why It Matters

Enables reliable character consistency across comic strips, storyboards, and visual stories — a game-changer for narrative AI workflows.

📬 Get the top 10 AI stories daily