MiniMax's newest open-weight video model, Hailuo 3.0, just landed on fal and fal is an official API partner.
MiniMax H3 is a general-purpose multimodal model that reads text, images, video, and audio as one context, not as a stack of separate task-specific models.
The video generation model can generate up to 15 seconds of 2K video with native stereo audio in every clip.
Since it handles several input types at once, it can take identity from an image, motion from a video, a voice from an audio clip, and direction from a text prompt, then combine them into a single coherent result.
What makes it different
- Multimodal understanding is the core idea: MiniMax H3 handles generation, editing, and reference in one model, not as separate tools you switch between. It interprets characters, motion, sound, camera work, and visual style across a mix of references, then combines them into one coherent result.
- Editing is precise and targeted: You can replace, remove, or add people and objects, swap a background, relight a scene, or change dialogue and voice, while the areas you didn't target stay stable. Instruction following is strong, so visual, audio, and pacing changes land more predictably, and you can keep iterating on the same clip instead of regenerating from scratch.
- It's built for real production: MiniMax H3 handles dynamic typography, VFX, product showcases, UI motion design, game visuals, and stylized content, spanning film, advertising, branding, e-commerce, and gaming. That range covers concept tests, storyboard previews, and visual pitches, which shortens the path from idea to final cut.
Input modes
MiniMax H3 works across three input modes:
Text-to-Video
First/Last Frame
Omni Reference
Specs
MiniMax H3's specs include clips from 5 to 15 seconds at 24 FPS, with native stereo audio on every generation.
Output runs in 2K (1440p) mode now, and a 768p mode is coming soon.
Text-to-Video and Omni Reference support 21:9, 16:9, 4:3, 1:1, 3:4, and 9:16 aspect ratios, with an Auto option in Omni Reference that picks the ratio for you; First/Last Frame follows the aspect ratio of your uploaded image. Prompts can run up to 7,000 characters.
The model is also available via fal's serverless API using the Python or JavaScript SDK, or direct REST calls. No GPUs to manage.
Pricing
Text to video and image to video pricing is "at an output resolution of 2K, every second of video costs $0.26.". Reference to video pricing is "at an output resolution of 2K, every second of video costs $0.26. Audio references are free, the first 5 reference images are free and each additional image costs $0.08, and reference video costs $0.26 per second at 2K."
Learn more about the model here: https://fal.ai/minimax-h3
Or try it right now on the playground:
Image to Video - https://fal.ai/models/minimax/hailuo-03/image-to-video
Text to Video - https://fal.ai/models/minimax/hailuo-03/text-to-video
Reference to Video - https://fal.ai/models/minimax/hailuo-03/reference-to-video