Video model · MiniMax
MiniMax H3
MiniMax's open-weights multimodal video model with native sound
What is MiniMax H3?
MiniMax H3 is a general-purpose multimodal video model that generates 5 to 15 second clips at up to 4K and 24 fps, with stereo audio on every generation. Text, images, video, and audio all share one context, so a single request can carry a character from a photo, the camera language from a clip, and a voice from a recording — up to nine reference images, three video clips, and three audio clips at once.
Capabilities
- Generate video from a text prompt
- Animate a still image into video
- Keep characters and products consistent with reference images
MiniMax H3 on Melius
Melius runs MiniMax H3 alongside every other major image, video, audio, and text model, so you can generate with it, compare it against the field, and let the Mel agent route to it automatically — all on one infinite canvas. You only pay for what you generate.
Frequently asked questions
What is MiniMax H3?
MiniMax H3 is a general-purpose multimodal video model that generates 5 to 15 second clips at up to 4K and 24 fps, with stereo audio on every generation. Text, images, video, and audio all share one context, so a single request can carry a character from a photo, the camera language from a clip, and a voice from a recording — up to nine reference images, three video clips, and three audio clips at once.
How do I use MiniMax H3 on Melius?
Melius is a node-based canvas: add a node, pick MiniMax H3 from the model picker, and connect it to whatever should feed it or follow it. You can also just describe what you want and let the Mel agent wire it up. You only pay for what you generate.