Research & Papers

Wan-Streamer v0.1: Real-time audio-visual AI with 200ms latency

End-to-end model handles video, audio, and text in a single Transformer.

Deep Dive

Wan-Streamer v0.1 redefines real-time multimodal interaction by eliminating the traditional cascade of separate modules (VAD, ASR, language model, TTS, avatar animation, video generation). Instead, it uses a single Transformer with interleaved visual, audio, and text tokens, coordinated by block-causal attention for incremental streaming. This design reduces pipeline latency and error accumulation, achieving native full-duplex audio-visual communication.

The model processes streaming units as short as 160 ms at 25 fps, with model-side response latency of approximately 200 ms and total interaction latency around 550 ms when combined with 350 ms bidirectional network latency. The system learns to handle perception, reasoning, generation, response timing, turn management, and cross-modal synchronization end-to-end. This positions Wan-Streamer as a unified foundation model for low-latency, interactive video and voice applications.

Key Points
  • ~200ms model latency and ~550ms total interaction latency (including network) for sub-second duplex communication.
  • Streaming units as short as 160ms at 25fps using causal encoders, decoders, and block-causal attention.
  • Single Transformer jointly models language, audio, and video input/output without external modules like VAD, ASR, TTS, or video generation.

Why It Matters

Wan-Streamer could enable truly natural, real-time video conversations with AI assistants, replacing clunky pipeline approaches.

📬 Get the top 10 AI stories daily