JoyAI-Echo video model on HF delivers minute-long stories 7.5x faster
46GB model generates coherent multi-shot stories with synchronized audio from one prompt.
Deep Dive
jd released a new video model based on LTX-2 (46GB) on GitHub, focused on long-form video. Key features: minute-level multi-shot stories from one prompt JSON, DMD-distilled few-step inference (~7.5x faster), joint audio-video generation, and a paired cross-modal memory bank for story-level consistency. The quality isn't groundbreaking, but may have fun or useful applications.
Key Points
- Minute-level multi-shot stories: generates a coherent sequence of shots from one JSON prompt.
- DMD-distilled few-step inference: ~7.5x faster than the original pipeline, reducing generation time significantly.
- Joint audio-video generation: a single pipeline produces synchronized video and audio, with a cross-modal memory bank ensuring story-level consistency.
Why It Matters
Enables rapid, long-form video and audio generation from one prompt, useful for storyboarding and prototyping.