MiniMax unveils H3 omnimodal model, says 2K video generation costs one-third of flagship peers
MiniMax on July 31 introduced H3, also branded as Hailuo 3.0, a new omnimodal generation model that the company said will have open weights released in the coming days. The launch follows MiniMax’s open-source release of its M3 text model in June and positions H3 as a system built to handle mixed text, image, video, and audio inputs in a single workflow rather than splitting those functions across separate models. According to the company, H3 can take text, images, audio, and video as inputs, including mixed prompts, and produce native 2K, 24 fps video clips running from 5 to 15 seconds with synchronized sound generated in the same pass. MiniMax said the price for 2K video generation is RMB 0.8 per second, or about RMB 12 for a 15-second clip, which it described as roughly one-third the cost of comparable flagship models. MiniMax also highlighted its Omni-Reference feature, which supports up to nine reference images, three reference videos, and three reference audio clips in a single generation. The model also supports text-to-video, image-to-video, multi-shot storytelling, instruction-based editing, and V2V Motion Transfer. Still, its broader standing remains unsettled: MiniMax’s claimed top ranking applies to video editing rather than overall video generation, while open weights have not yet appeared on Hugging Face.








