MiniMax unveils Hailuo 3.0 (H3) video model with native 2K and synced audio
The model generates synchronised dialogue, sound effects and ambient audio alongside native 2K video in a single pass, and accepts up to nine reference images for consistency.
- Models & capabilities
- Minor
Chinese AI company MiniMax unveiled Hailuo 3.0, also called H3, at the World Artificial Intelligence Conference (WAIC) in Shanghai on 17 July 2026, ahead of a public release two weeks later. The video-generation model produces native 2K clips at 24 frames per second, running from five to fifteen seconds and extendable to roughly thirty, with synchronised dialogue, sound effects and ambient audio generated in the same pass rather than added afterwards.
The model supports “omni-reference” control — up to nine reference images, three video clips and three audio clips to hold characters, settings or voices consistent across generations — plus instruction-based editing that lets users describe a change to an existing clip rather than regenerate it from scratch. MiniMax pitched H3 on control and editing speed rather than raw resolution: rival Kling 3.0 offered native 4K output, and Google’s Veo 3.1 was seen as stronger on dialogue fidelity. Independent benchmark comparisons were not yet available at unveiling.
The release continued a period of rapid iteration among Chinese video-generation labs, arriving several months after OpenAI discontinued its Sora web and app products, leaving less US competition in the consumer video-generation space that MiniMax, Kling and others were contesting.