What is MiniMax H3?
MiniMax H3 is a new open-weight video model that generates footage, dialogue, and sound together in a single pass. This page explains what MiniMax H3 does, how its modes work, and what its specs mean in practice.
One model, one context
Most AI video tools are really a chain of separate models stitched together: one system writes a script, another turns text into an image, a third animates that image, a fourth generates a soundtrack, and a fifth syncs voice to lips. Each handoff loses information, and each stage has its own quirks, its own failure modes, its own prompt syntax. MiniMax H3 model takes a different approach. Text, images, video references, and audio all sit inside one unified context, processed by a single model rather than routed between task-specific endpoints. When you ask what is MiniMax H3 at its core, this is the honest answer: it is an omni-modal system that reasons about a shot the way a director does, holding the subject, the camera move, the lighting, and the soundtrack in mind at the same time, instead of solving each piece in isolation. That unified context is why a MiniMax H3 clip feels coherent rather than assembled, since the dialogue lands on the cut, the score matches the mood of the lighting, and a character's voice stays consistent across a reference and a new performance, because nothing was generated in a separate silo.
Four ways to start a shot
MiniMax H3 supports four modes, and each answers a different question about how a shot begins. Text-to-video starts from nothing but a written prompt, letting the model invent framing, motion, and sound from a description alone. First-and-last-frame mode takes two stills, where the shot opens and where it lands, and generates the motion, lighting shift, and audio that connects them, which is useful when the beginning and end of an action already exist and only the middle needs to be built. Reference-to-video leans on the model's omni-modal context directly: feed it images for a subject or style, clips for a motion pattern, or audio for a voice or score, and it extends those references into a new, longer shot. Natural-language editing treats an existing clip as a draft rather than a final cut, letting you swap a product on a shelf, rewrite the text on a sign, relight a scene from day to night, or add and remove an object, all through a plain-language instruction that changes only the part described and leaves the rest of the shot stable.
Resolution, duration, and frame rate
Every MiniMax H3 generation renders at 2K by default, which in practice means 1440 pixels on the short edge across every supported ratio between 16:9 and 9:16, and at the wide end, 21:9, that works out to roughly 3.7 megapixels, or 2976x1248. That is enough detail to hold up on a large screen without the soft, upscaled look common to fast video generators. Clips run between 5 and 15 seconds, rendered at a cinematic 24 frames per second rather than a choppier or artificially smoothed rate. Seven aspect ratios are available, 21:9, 16:9, 4:3, 1:1, 3:4, and 9:16, plus an adaptive mode that picks a ratio to match the composition implied by the prompt or references, useful when you are not sure in advance whether a shot reads better wide or tall. Together these numbers describe a model built around a specific kind of output: short, high-resolution, well-paced clips suited to social platforms, product spots, and short-form storytelling, rather than long-form footage where duration matters more than per-frame fidelity.
Sound is not an afterthought
Every MiniMax H3 generation ships with native stereo audio, not a silent clip waiting for a separate sound pass. The model generates an original score, dialogue, foley, and room tone together with the picture, timed to match the cut rather than layered on afterward, so a door closing lands on the frame where it visibly closes, and a musical swell arrives with the visual beat it is supporting. This matters because sound design usually happens after picture lock in most workflows, and that gap is where mismatched timing creeps in. MiniMax H3 also supports voice transfer: give it a reference recording, and it can carry that voice, its tone, pacing, and character, onto a performance in the generated scene, rather than defaulting to a generic synthetic narrator. That makes it possible to keep a consistent voice across multiple generations of the same character, or to have a single recorded line of dialogue drive a new scene entirely. The result is a clip that sounds designed, not dubbed, straight out of a single generation.
Spending the reference budget
MiniMax H3 accepts up to 12 reference files in a single generation: as many as 9 images for subject or style, up to 3 video clips totaling 15 seconds for motion, and up to 3 audio tracks totaling 15 seconds for voice or score, all inside the same unified context described above. The images are the most flexible slot: use them for a character's face, a product's exact shape, or a location's color palette, and combine several to lock in both a subject and a style at once. Video references are best spent on motion you cannot easily describe in words, a specific camera move, a walk cycle, a gesture, since the model reads them for movement rather than appearance. Audio references carry a voice or a musical identity forward into a new shot. The budget rewards restraint: a handful of well-chosen references that each do one clear job, one for the face, one for the camera move, one for the voice, outperforms stuffing all 12 slots with redundant material that gives the model conflicting signals to reconcile.
What open weights actually changes
MiniMax H3 launched with open weights, which means the model itself, not just an API in front of it, is available for anyone to inspect, run, and build on. That is a meaningfully different posture from a closed model reachable only through a paid endpoint whose behavior can shift without notice. MiniMax H3 open weights means researchers can study exactly how the omni-modal context is put together, developers can host the model on their own infrastructure instead of depending on one vendor's uptime, and the wider community can fine-tune it for specialized use cases the original release never anticipated. It also means the pace of improvement is not gated behind a single company's roadmap, since anyone can contribute fixes, adapters, or optimizations. For a model this capable, open weights is the detail that turns a single product launch into a shared piece of infrastructure that a whole ecosystem of tools, tutorials, and derivative models can grow around, rather than a walled garden that only one company controls the future of.
Writing a prompt MiniMax H3 can act on
MiniMax H3 accepts prompts up to 7,000 characters, which is enough room to direct a shot the way you would brief a small crew rather than typing a caption. Start with the subject: who or what is in frame, and what they are doing. Add the camera: is it a slow push-in, a static wide shot, a handheld follow, since naming the move gives the model something concrete to execute rather than guessing. Describe the lighting and time of day, since that detail drives both the visual mood and, indirectly, how the model paces the shot. Because audio generates alongside picture, describe the soundscape too, a quiet room tone, a specific line of dialogue, a swell of score at a particular beat, instead of leaving sound as an afterthought in the text. Finally, give a sense of editing rhythm: does the shot hold steady, or build to a cut. A prompt that covers all five, subject, camera, lighting, soundscape, and rhythm, gives MiniMax H3 a complete brief instead of a fragment it has to fill in on its own.
Specifications
MiniMax H3 at a glance
| Feature | MiniMax H3 |
|---|---|
| Resolution | 2K (2976x1248 at 21:9) |
| Duration | 5-15 seconds at 24 FPS |
| Audio | Native stereo, every generation |
| Aspect ratios | 7 ratios plus adaptive |
| References | 9 images, 3 clips, 3 audio (12 max) |
| Prompt length | Up to 7,000 characters |
| Modes | Text, first/last frame, reference, editing |
| Weights | Open |
Try it
Try MiniMax H3 right here
This first step renders the opening frame of your shot as a still image, so you can check framing and lighting before committing to the full 2K clip with audio, a quick, free way to preview composition before the complete generation runs.
Free and anonymous. Protected by Cloudflare Turnstile — no account, no card.
FAQ
MiniMax H3 questions answered
MiniMax H3 is an open-weight, omni-modal AI video model released July 31, 2026. Instead of routing text, image, video, and audio requests through separate task-specific systems, MiniMax H3 processes all of them inside one unified context, generating a full shot, picture, motion, dialogue, and score, from a single request. It outputs 2K clips of 5 to 15 seconds at 24 FPS across seven aspect ratios, with native stereo audio and voice transfer built into every generation, not added as a separate step afterward.
Open weights means the underlying MiniMax H3 model, not just a hosted API, is publicly available to download, inspect, and run. That lets developers self-host the model instead of depending on one company's servers, lets researchers study how its omni-modal context actually works, and lets the community fine-tune or extend it for use cases the original release never covered. It is the difference between a closed product you can only rent and shared infrastructure a whole ecosystem can build on.
A single MiniMax H3 generation runs 5 to 15 seconds at 24 FPS, rendered at 2K, 1440 pixels on the short edge, up to 2976x1248 at the widest 21:9 ratio. Every clip includes native stereo audio by default: an original score, dialogue, foley, and room tone generated together with the picture and timed to the cut, rather than added in a separate pass afterward. Voice transfer can also carry a reference recording's voice onto a character in the new scene.
MiniMax H3 accepts up to 12 reference files per generation: up to 9 images for a subject or visual style, up to 3 video clips totaling 15 seconds for a motion pattern, and up to 3 audio tracks totaling 15 seconds for a voice or score. All of them sit inside the same omni-modal context the model reasons over, so a face from one image, a camera move from one clip, and a voice from one track can combine into a single new shot.
Explore
See what MiniMax H3 can generate
MiniMax H3 examples are the fastest way to understand what this model actually produces, real clips with the exact MiniMax H3 prompts that made them, ready to reuse. Browse the gallery to see the four modes, the native audio, and the reference workflow in action, then try the tool above on a shot of your own.
Ready to generate 2K video?
Experience the unified context, native audio, and open-weights freedom of MiniMax H3 today.