MiniMax H3: A Practical Guide to Open-Weight Multimodal Video Generation
MiniMax H3 combines text, images, video, and audio to generate short videos with native stereo sound. Here is what the official model card confirms, what local deployment does and does not include, and how teams should evaluate it.

Official MiniMax H3 resources
- MiniMax H3 official announcement
- MiniMax H3 model card and deployment guide
- MiniMax H3 Community License
- MiniMax video generation API reference
This article uses these first-party documents as the primary sources. It does not embed or rely on third-party social-media demonstrations.
The short version
MiniMax H3 is an open-weight, omni-modal video generation system designed to combine text, images, video, and audio in one request. The official model card specifies 4 to 15 second video output, native stereo audio, supported 768P and 2K workflows, and a reference mode with images, videos, and audio.
The important distinction is operational. MiniMax has released H3 model weights, but its full official 2K pipeline includes hosted components. A team considering local deployment should separate what can run locally today from what depends on the hosted workflow and license conditions.
What MiniMax H3 officially supports
| Area | Officially documented capability |
|---|---|
| Video duration | 4 to 15 seconds |
| Output | Video with 32 kHz stereo audio |
| Resolution | 768P locally through H3-Base, with a 2K regeneration workflow documented by MiniMax |
| Aspect ratios | Includes 21:9, 16:9, 4:3, 1:1, 3:4, and 9:16 |
| Dialogue languages | 11 languages, including Japanese, English, Chinese, Korean, French, German, Italian, Portuguese, Russian, Spanish, and Arabic |
| Generation modes | Text to audio-video, first or last frame, first and last frame, and multimodal reference to video |
| Reference inputs | Up to 9 images, 3 videos, and 3 audio clips in the reference mode, with 12 files total |
MiniMax describes H3 as a general-purpose omni-modal generation model. In practice, that means the prompt can state the relationship among references instead of treating image, motion, and audio as disconnected jobs.
H3's prompt model: describe relationships among media
The official announcement uses a concise example: use camera movement from a reference video, a character from a reference image, and vocals from a reference audio clip. This is a useful way to think about H3 prompts.
Use Video 1 for the slow circular camera movement.
Use Image 1 for the presenter’s appearance and clothing.
Use Audio 1 for the vocal timbre.
The presenter walks through a quiet gallery, stops beside the product pedestal,
and says the supplied Japanese line in a calm, natural voice.
Keep the gallery lighting and the product shape consistent.
The goal is not to upload the largest possible collection of files. It is to specify which reference governs identity, motion, scene, sound, or visual style. If a reference is ambiguous, state what must not be inherited as well.
Three modes, three planning patterns
Text to audio-video
Use this when the scene, subject, motion, and audio can be described without a source asset. Text-to-video requires a non-empty text prompt and an explicit output ratio.
First and last frame
Use a first frame to establish the opening composition. Add a last frame when the end composition is important. This mode is useful for product transformations, logo-to-scene transitions, and planned opening or closing shots.
Omni reference
Use reference mode when identity, product geometry, motion, or audio needs to follow supplied materials. Reference audio cannot be the only source asset. The API requires at least one reference image or reference video alongside it.
What is open-weight, and what remains hosted
The model card describes three parts of the broader system:
- H3-Context-IR interprets free-form multimodal inputs and turns them into a structured intermediate representation.
- H3-Base generates audio-video at 768P.
- H3-Regenerate-2K uses the original context and a lower-resolution result to create a 2K version.
The initial open release includes H3-Base checkpoints for local inference. MiniMax states that Context-IR and Regenerate-2K are hosted components, with an API-supported workflow for validating official 2K results. This matters for planning. “Open weights” does not mean every production component can be self-hosted in the same way today.
A practical local evaluation plan
Before committing an internal workflow or a campaign to H3, run a controlled test rather than selecting from social clips.
- Choose one product, one character, one 5 to 10 second action, and one output ratio.
- Run the same brief in text-only, frame-conditioned, and reference modes where appropriate.
- Record resolution, duration, hardware, elapsed time, retries, prompt version, and every input asset.
- Review identity, product shape, motion, audio synchronization, dialogue, and text accuracy against a checklist.
- Separate a promising visual result from an approved production asset. Add exact brand text and legal copy in finishing when needed.
This test gives teams a usable baseline for cost per approved output, not merely an impressive single example.
Official examples and a safe evaluation checklist
For this guide, we are not embedding third-party creator demos. The relevant public examples are MiniMax's own product demonstrations, which cover film opening titles, product websites, animated posters, advertising, and e-commerce. View the official H3 examples in MiniMax’s announcement.
When evaluating H3 with your own assets, use a short review checklist:
| Review area | What to check |
|---|---|
| Reference mapping | Is it clear which input controls identity, product appearance, motion, scene, or voice? |
| Visual continuity | Do subject identity, product geometry, screen direction, and lighting remain stable? |
| Audio | Are dialogue, timing, ambience, and stereo sound appropriate for the action? |
| Brand accuracy | Are wordmarks, UI labels, and regulated copy added from approved source assets where needed? |
| Operations | Are hardware, duration, resolution, retries, and review time recorded? |
This keeps the evaluation tied to a team’s own brief, permissions, and brand standards rather than a creator clip selected from social media.
License and deployment points to verify
The MiniMax H3 Community License deserves a separate review before a commercial rollout.
- The listed applicable territory excludes the European Union, United Kingdom, Republic of Korea, and United States unless a separate license is obtained.
- Commercial products or services above US$20 million in annual revenue require prior written authorization from MiniMax.
- Commercial interfaces using H3 must prominently display “MiniMax H3.”
- The license requires safeguards for products or hosted services that let others generate outputs.
- The license restricts using H3 works or outputs to improve another AI model.
This is a summary, not legal advice. Review the current license text and the applicable laws with counsel before deployment.
H3 and Seedance 2.5: choose based on the workflow
Seedance 2.5 and MiniMax H3 were announced in the same release window, but their operational emphasis differs.
| Question | MiniMax H3 | Seedance 2.5 |
|---|---|---|
| Native output duration stated by the provider | 4 to 15 seconds | Up to 30 seconds |
| Model availability direction | Open weights with explicit license conditions | Product and platform rollout |
| Reference workflow | Text, image, video, and audio in omni-reference mode | Larger multi-reference production brief and staged long-form prompting |
| 2K workflow | Documented through base generation plus regeneration workflow | Provider product workflow |
| Best question to test | Can we run and govern an open-weight workflow under the license? | Can we direct a longer multi-stage sequence with our assets? |
Do not infer a winner from a single demo. Choose a brief, run matched tests, record the review criteria, and compare the output that your team can approve and use.
Sources and evidence boundaries
- MiniMax official announcement, July 31, 2026. Capability and product-positioning claims.
- MiniMax H3 model card. Input limits, output specifications, workflow modules, and deployment guidance.
- MiniMax video API reference. Current request structures, supported roles, media constraints, aspect ratios, duration, and asynchronous task flow.
- MiniMax H3 Community License, dated August 2, 2026. Territory, commercial, attribution, and safeguard conditions.
Turn a model test into a shared production workflow with Kylon
A video model test becomes more useful when the materials, decisions, and review criteria stay with the team rather than one creator’s browser session. Kylon gives a creative team one shared place for the brief, approved inputs, test outputs, review notes, and next actions.
Keep a reusable brand-material pack
Create a shared brand kit for each video workflow: product imagery, approved logo files, UI captures, color and typography guidance, terminology, voice direction, mandatory copy, and usage boundaries. The team can attach those materials to the project conversation and keep versioned source assets alongside the brief.
For H3, that means a reference image is not just an upload. The team can record what it is approved to define, such as product geometry, wardrobe, environment, or camera composition, plus what must not be inherited.
Share the production skill, not only the final prompt
A Kylon Skill can hold the repeatable method behind a successful test: how to prepare source files, map multimodal references, write a prompt, check output quality, add approved brand text, and route the result for review. Once shared, the same skill is available to the team instead of remaining in one person’s notes.
Review with the full context visible
Use a Kylon channel to keep the brief, source files, generated versions, comments, and approval decision together. Agents can help prepare a first draft of the prompt or organize a comparison, while people retain the approval step for brand-sensitive materials and external publishing.
If your team is evaluating MiniMax H3, Seedance, or another video model, Kylon can provide the shared workflow around the model: brand-material governance, reusable production skills, review context, and a record of what was approved.
Bottom line
MiniMax H3 is significant because it brings a multimodal video system and an open-weight deployment path into the same conversation. Its practical value depends on the exact workflow: a clear role for every input, measured local or API tests, human review of audio and visuals, and careful attention to the current license.
Hire the AI agent team that runs your entire business.

