Google's Gemini Omni unifies video, audio, and text in one model
One model takes any input—text, image, audio, video—and outputs clips, edits, or avatars.
At Google I/O 2026, Google unveiled Gemini Omni—a unified multimodal AI model that replaces the previous stack of separate models (Veo 3.1, Imagen 3, Lyria) with a single system capable of reasoning across text, image, audio, and video inputs. CEO Sundar Pichai framed it as "create anything from any input." The model outputs video clips, edited photos, or custom digital avatars, all with shared context across modalities. Gemini Omni Flash launches today, supporting 4–10 second clips at up to 1080p, with direct integration into the Gemini app and YouTube Shorts. API access will follow in the coming weeks.
Key features include conversational editing—users can refine clips via chat without restarting—and long-context consistency that preserves character identities across shots. Google demonstrated physics-aware generation (e.g., a marble rolling with realistic gravity and synchronized audio) and scene-aware edits (e.g., turning a sculpture into bubbles). All outputs carry SynthID watermarking. Custom avatars are created by recording a number sequence, avoiding real-person generation. A more capable Gemini Omni Pro was teased without a release date. The move signals Google's shift from standalone generative video tools to a core Gemini capability, making Omni the new center of gravity for multimodal creation.
- Gemini Omni unifies text, image, audio, and video in one model—no more chaining separate tools like Veo or Imagen.
- Flash model is live today with 10-second 1080p clips, conversational editing, and SynthID watermarking.
- Pro version teased; API access coming in weeks. Custom avatars via number sequence recording.
Why It Matters
Google collapses the entire generative video workflow into one model, enabling seamless cross-modal creation and editing.