Open Source

MiniMax H3 generates 2K video with stereo sound, open weights coming

15-second 2K clips with native audio at under 1/3 the price of rivals

Deep Dive

MiniMax unveiled H3, a general-purpose multimodal generation model designed for commercial content creation. H3 unifies context across text, images, video, and audio, enabling generation of up to 15 seconds of 2K-resolution video with native stereo sound. Early tests highlight strong instruction following, accurate text and brand rendering, and V2V (video-to-video) motion transfer. The model uses a stack of proprietary technologies — Contextual Omni Representation, H3-VAE, H3-Omni Transformer, and In-Context Regeneration — to deliver what the company calls industry-leading price-performance. At 2K resolution, H3's per-second price is less than a third of mainstream models; at 768p, it falls below half the cost of competitors' 720p output.

MiniMax positions H3 for advertising, branding, e-commerce, product design, UI/UX, and gaming, with precise controllable generation and editing. The company also criticizes closed-source video models for slower iteration and a less open ecosystem, so it plans to release H3's model weights in the coming days, subject to applicable laws. Hardware compatibility was considered from the earliest design stages, aiming to accelerate support across a broader range of AI accelerators. This move could let developers fine-tune and self-host H3, potentially disrupting the closed video-generation market and lowering costs for high-fidelity, audio-synced video production.

Key Points
  • MiniMax H3 outputs up to 15 seconds of 2K video with native stereo sound
  • Pricing undercuts mainstream models: 2K costs less than 1/3, 768p less than half
  • Open weights planned within days, with early hardware compatibility across AI accelerators

Why It Matters

Open-weight video generation at this quality could lower content production costs and break closed-model dominance.

📬 Get the top 10 AI stories daily