MiniMax pushed the weights for H3, its audio-video generation model, to Hugging Face in the early hours of August 3. The repository's own commit history dates the initial content commit to roughly 02:30 UTC and the diffusers-format weights to about 04:30 UTC. The model was announced on July 31 with a promise to open the weights "in the coming days."
What shipped
H3 is a 33-billion-parameter dense single-stream transformer, roughly 13 billion of it sitting in AdaLN branches that can be cached out of the inference loop. It produces 4 to 15 seconds of video at 24 FPS, 768p by default and up to 2K, with native 32 kHz stereo audio across 11 languages. Its text encoder is built on Qwen3-VL-32B.
The licence
The bundled licence is not an open-source licence and Hugging Face tags it simply as other. It reads: "You may not use, reproduce, modify, distribute, or display the MiniMax H3 Works or any of their Outputs or results outside the Applicable Territory." Applicable Territory means worldwide minus the EU, the UK, South Korea and the United States. Two further clauses bar using H3's outputs to improve any other AI model, and require written authorisation above $20 million in annual revenue.
Where the framing breaks
"MiniMax open-sources H3" fails twice over. A reader in London, Seoul, Berlin or San Francisco is not licensed to run it at all — the export restriction points the other way for once. And the capability everyone quotes, 15-second 2K video with synchronised audio, needs the two modules that stayed closed.
Three hours to a domestic port
Moore Threads said it carried H3 from the SGLang-MUSA framework through its MATE and muDNN operator libraries on its MTT S5000 in three hours, the same day. That turnaround, not the model, is the harder signal.
