This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.

MiniMax H3 Opens Weights and Overtakes Seedance 2.0 on Key Video Benchmarks
N
News EditorMiniMax released the weights for its omni-modal video model H3 on August 3, posting them on Hugging Face. The model can generate 2K, 24fps video with stereo audio, and it moved ahead of ByteDance’s Seedance 2.0 in the Artificial Analysis text-to-video ranking while taking first place in video editing. Pricing also came in lower: a 5-second clip costs about $0.70 for H3’s 2K output versus roughly $1.22 for Seedance 2.0’s 720p output. Community tests also showed the model running offline on consumer GPUs with 12GB of VRAM. The open-weight release, however, stops at 768p; 2K output still depends on a non-open API-side upsampling module, and the Community License carries regional and usage limits.
MiniMax put the weights for its H3 omni-modal video model on Hugging Face on August 3, and the release quickly made its way into ComfyUI and other tools. H3 can generate 2K, 24fps video together with stereo audio, and the model has already been used in open workflows, third-party quantized builds and API aggregators. MiniMax first introduced H3 through an API on July 31 before opening the weights three days later.
According to the official model card, H3 is a 33B-parameter dense Transformer. About 13B of those parameters sit in an AdaLN modulation branch that does not need to be loaded at inference time. The model takes text, images, video and audio into one shared context and generates output in a single pass. It can produce clips of up to 15 seconds, with 24fps video and 32kHz stereo audio created at the same time. Dialogue, sound effects, ambient audio and visuals are generated together rather than stitched in later.
H3 supports up to 9 reference images, 3 reference video segments and 3 reference audio segments in a single run. It also covers text-to-video with audio (T2VA), first-last frame control (FL2VA) and reference-based generation (Ref2VA). The model spans six aspect ratios, from 21:9 to 9:16, and supports stable dialogue in 11 languages. The open-weight release is the 768p H3-Base model; the 2K output is produced through the API-side H3-Regenerate-2K upsampling module.
The benchmark picture is mixed, but it still tilts toward H3 in key places. In Artificial Analysis’ early-August rankings, H3 took first place in video editing. In text-to-video, it ranked above ByteDance’s Seedance 2.0, while Seedance 2.0 remained first in image-to-video. Seedance 2.0, released by ByteDance in February, also focuses on joint audio-video generation and multi-shot character consistency, and later updates added native 4K output and 10-bit color depth.
Pricing is where the gap becomes even clearer. For a 5-second clip, H3’s 2K output costs about $0.70, while Seedance 2.0’s 720p output costs about $1.22 under its token-based pricing. On that basis, H3 comes in at roughly 57% of the cost. At the 1080p level, third-party estimates put H3 at about $7.8 per minute, compared with about $22.45 for Seedance 2.0.
Availability is another dividing line. Seedance 2.0 was briefly halted in global release because of Hollywood copyright issues, and Seedance 2.5, released in June, remains closed and tied to ByteDance’s own platforms. H3 weights are already downloadable from Hugging Face, where developers can fine-tune them today. The catch is that H3 uses MiniMax H3 Community License rather than a standard open-source license, and the terms place restrictions on certain regions and use cases.
The ecosystem around H3 has already started to form. MiniMax showed a game-cutscene demo that generated an ancient water temple, a glowing artifact, giant stone guardians and an escape sequence, all with synchronized audio. fal joined as a day-one partner, and Luma added H3 to its multi-model Agents product on August 6. fal published 44 generation examples covering product ads, character storytelling and interface animation, while ComfyUI documentation added workflows for text-to-video, image-to-video and reference generation.
Community use cases are moving quickly into production-style workflows. WeShop, an e-commerce tool, showed a setup built around a single product image and one viral reference video. H3 then transferred the pacing, shot structure and sales style of the reference clip onto the merchant’s own product, producing short videos that can be pushed to TikTok and Reels.
Lip-sync workflows changed as well. What used to require TTS, a lip-sync model and post-production mixing can now be produced in one API call, because the mouth movement and voice are created together in the same generation pass. Quantized community builds have reduced the weights to 42.5GB, and ComfyUI said consumer GPUs with 12GB of VRAM can run the model after weight offloading.
That said, H3 is not fully open without limits. The open-weight release tops out at 768p, with the headline 2K output still relying on the non-open API upsampling module. Local runs of higher-quality versions still depend on quantization and weight-offloading tricks. And the Community License adds region and use restrictions that developers need to read closely before any commercial deployment.
190
Disclaimer:
The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.
Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.
