MiniMax unveils H3 omnimodal model, says 2K video generation costs one-third of flagship peers

MiniMax unveils H3 omnimodal model, says 2K video generation costs one-third of flagship peers

N
News Editor
2026-07-31 03:52:34
MiniMax on July 31 introduced H3, also branded as Hailuo 3.0, a new omnimodal generation model that the company said will have open weights released in the coming days. The launch follows MiniMax’s open-source release of its M3 text model in June and positions H3 as a system built to handle mixed text, image, video, and audio inputs in a single workflow rather than splitting those functions across separate models. According to the company, H3 can take text, images, audio, and video as inputs, including mixed prompts, and produce native 2K, 24 fps video clips running from 5 to 15 seconds with synchronized sound generated in the same pass. MiniMax said the price for 2K video generation is RMB 0.8 per second, or about RMB 12 for a 15-second clip, which it described as roughly one-third the cost of comparable flagship models. MiniMax also highlighted its Omni-Reference feature, which supports up to nine reference images, three reference videos, and three reference audio clips in a single generation. The model also supports text-to-video, image-to-video, multi-shot storytelling, instruction-based editing, and V2V Motion Transfer. Still, its broader standing remains unsettled: MiniMax’s claimed top ranking applies to video editing rather than overall video generation, while open weights have not yet appeared on Hugging Face.

MiniMax on July 31 announced H3, also called Hailuo 3.0, a new omnimodal generation model, and said the model weights will be open-sourced in the coming days. The release comes after the company open-sourced its M3 text model in June.

The company described H3 as a move away from task-specific models and toward general multimodal intelligence. Instead of separating image generation, video generation, audio generation, editing, and reference handling into different tools, H3 is designed to interpret what a user wants inside a mixed context of text, images, video, and sound.

MiniMax says 2K video costs RMB 0.8 per second

MiniMax said H3 can generate 2K video at RMB 0.8 per second. A 15-second clip would cost about RMB 12, according to the company, which said that works out to about one-third the cost of comparable flagship models.

The company’s official account also posted that H3 is available on Runway and added that open weights are coming soon.

Mixed multimodal input and native synchronized output

On the input side, H3 accepts text, images, audio, and video, and those inputs can be combined in a single prompt. On the output side, it produces native 2K, 24 fps clips ranging from 5 to 15 seconds with synchronized sound. MiniMax said the visuals and audio are generated together in one pass, rather than creating video first and adding sound later.

A central feature is what MiniMax calls Omni-Reference. In one generation, users can include as many as nine reference images, three reference videos, and three reference audio clips to specify style, character, motion, or voice characteristics. MiniMax said those references are not treated as key frames that must be copied directly, leaving the model room for its own interpretation.

Video generation, editing, and motion transfer

MiniMax said H3 supports text-to-video, image-to-video, multi-shot storytelling, instruction-based editing, and V2V Motion Transfer, which refers to transferring motion from one video clip to a different character.

For use cases, the company pointed to advertising, brand marketing, e-commerce, product design, UI/UX, and gaming. Its wording emphasized “commercial-grade production” rather than simply video generation.

Performance claims and open-source timeline still need outside verification

MiniMax’s claimed first-place result applies to the video editing subcategory, not to overall video generation.

On the same Artificial Analysis ranking, data through July showed ByteDance’s Seedance 2.0, released in February this year, still leading the video-with-audio category with 1,213 Elo. In image-to-video blind testing, it ranked first with 1,344 Elo. Google Veo 3.1 was described as the strongest overall performer, while Kuaishou’s Kling 3.0 led the llm-stats video arena with a score of 2,023.

As for open source, MiniMax has so far only said the release will happen “in the coming days,” and the weights had not appeared on Hugging Face at the time referenced in the source material. Whether H3 can support the company’s broader positioning will become clearer only after the weights are published and third-party benchmark results are available.

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
690

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.