MiniMax H3 is the strongest AI model for instruction-based video editing and multi-reference consistency, not the strongest for one-shot text-to-video — and treating those two jobs as the same thing is the mistake that quietly wastes the whole model. On the Artificial Analysis Video Arena (data through July 2026), H3 ranks #1 in Video Editing with Audio, but only #2 in Text-to-Video and #3 in Image-to-Video. If your job is keeping a character, product, or voice consistent across a multi-shot clip, or iterating on footage you already have, H3 is the right pick. If you just need one quick clip from a single line of text, a cheaper faster model is the better fit — H3's reference slots and editing controls become overkill, not magic.
For small teams and builders, the practical difference is this: most "AI video models" are really AI video generators — silent or audio-bolted-on, single-frame-in, clip-out. H3 is the small but growing category of AI video editing models — you bring finished footage (or a pile of references), describe what to change in plain words, and only the part you asked about changes while the rest stays pixel-stable. That distinction, more than the 2K resolution or the open-weights commitment, is what decides whether H3 belongs in your kit.
Last verified: August 2, 2026 (UTC)
- #1 Video Editing with Audio (Elo 1130) and #2 T2V with Audio (Elo 1242) on Artificial Analysis — leader in editing, top-3 elsewhere.
- 2K (2560×1440) output, 4–15 s per clip, native stereo audio — confirmed by MiniMax platform docs and the official blog.
- ~$0.13/sec ($7.80/min) at 2K with audio — undercuts Kling 3.0 1080p Pro ($20.16/min) and Veo 3.1 ($24.00/min) on price-per-minute at 2K.
- Open weights planned "in the coming days" under a MiniMax community license — not yet published; the API and consumer app are live today.
- Pricing, leaderboard Elo, and license terms are volatile facts — re-check before committing production work.
Where does MiniMax H3 rank against Veo, Kling, and Seedance?
H3 leads the field in video editing but is not the overall #1 video model — that nuance is the difference between "this model beat everything" (the headline) and "this model beat everything at editing" (the leaderboard). The third-party measurement comes from Artificial Analysis's Video Arena, which ranks models by blind human-preference votes on outputs from the same prompt — not vendor self-reporting.
Here is where H3 sits across the three boards that matter, with data current to July 31, 2026:
| Arena (With Audio) | #1 | #2 | #3 | H3 Elo | Best open-weights alternative |
|---|---|---|---|---|---|
| Video Editing | MiniMax H3 | Google Gemini Omni Flash | Alibaba-ATH HappyHorse-1.0 | 1130 | n/a (no open model ranked yet on this board) |
| Text-to-Video | Google Gemini Omni Flash | MiniMax H3 | Dreamina Seedance 2.0 720p | 1242 | LTX-2.3 Fast (Elo 980) |
| Image-to-Video | Dreamina Seedance 2.0 720p | Google Gemini Omni Flash | MiniMax H3 | 1184 | LTX-2.3 Pro (Elo 956) |
Three things to read from this table:
- Editing is the real story. H3 is the only model that places top-3 on all three boards, and its only leaderboard win is editing — a narrow, hard category most video models still treat as a separate pipeline. If a benchmark lead matters for your use case, this is the one.
- T2V and I2V go to Google and ByteDance, not MiniMax. H3 trails Gemini Omni Flash in T2V by 3 Elo points (1245 vs 1242, within the ±11 confidence interval) and trails both Seedance 2.0 and Gemini Omni Flash on I2V. The headline "beat the biggest models" is true where it counts most, not everywhere.
- Open-weights models sit well below H3 on every board. The best open model on T2V is LTX-2.3 Fast at Elo 980 — about 260 points behind H3. If MiniMax ships the open-weights release as planned, H3 would become the strongest open video model by a wide margin, displacing the LTX-2.3 family.
What can MiniMax H3 actually do that other models can't?
H3's differentiation lives in two places: omni-reference (one model reading image, video, and audio inputs together) and instruction-based editing (changing a finished clip without regenerating it). Veo, Kling, and Seedance will happily generate from a text prompt or a single image, but they don't take 12 reference files in one request and they don't accept "swap the sign, keep the rest" edits as a first-class operation.
Per the MiniMax platform documentation, one H3 request can include:
| Input type | Max per request | What it controls |
|---|---|---|
| Reference images | 9 | Character identity, product appearance, visual style, on-screen text/brand rendering |
| Reference video clips | 3 (2–15 s each, ≤15 s total) | Motion, choreography, camera movement, editing rhythm, performance |
| Reference audio clips | 3 (2–15 s each, ≤15 s total) | Voice timbre, sound design, music pattern — must be paired with an image or video, never sent alone |
| Total files | 12 | Combined cap across all reference types |
That's the technical answer to "how does H3 keep a character looking the same across shots?" — you feed it identity, motion, and voice as separate references and it carries all three into one finished output, instead of relying on a separate face-fix or voice-clone pass bolted on afterward.
Instruction-based editing: the capability ranked #1
The leaderboard H3 actually leads is video editing, not generation. What that means in practice, per MiniMax's API guide:
- You pass in a clip you already have.
- You describe the change in natural language ("swap the product on the desk for a black laptop," "relight the scene from day to night," "replace the line of dialogue at 0:08 with 'Welcome to the team'").
- The model edits only the portion you asked about — unedited regions stay pixel-stable, and the audio stays in sync.
That last point is the differentiator. Models like Kling 3.0 and Veo 3.1 will produce a strong new clip from scratch, but editing an existing one — keeping the parts you like and changing one thing — was historically a re-generation gamble or a manual editor task. H3 turns it into a one-call operation, and the Artificial Analysis Video Editing leaderboard places it ahead of Gemini Omni Flash (Elo 1121) and HappyHorse-1.0 (Elo 1096) on this exact capability.
For business teams that generate ads, product videos, or short social content and then need to A/B test ten variants of one winning clip (different on-screen text, different background, different closing line), editing the working footage is dramatically cheaper than ten fresh generations. That's the workflow H3 was built for.
How much does MiniMax H3 cost compared to Veo, Kling, and Seedance?
MiniMax H3 is priced at roughly $0.13 per second ($7.80/minute) for 2K output with native audio, per the Artificial Analysis pricing column and confirmed in MiniMax's official blog post which claims "at 2K, H3's per-second price is less than a third of mainstream models."
Compare the leading models on the Artificial Analysis boards, all at default 2K-or-equivalent with audio where available:
| Model | Creator | $/min | Resolution | Audio? | Source |
|---|---|---|---|---|---|
| MiniMax H3 | MiniMax | $7.80 | 2K (2560×1440) | Yes — native stereo | AA leaderboard |
| Gemini Omni Flash | $6.00 | (provider's default tier) | Yes | AA leaderboard | |
| Dreamina Seedance 2.0 720p | ByteDance | $9.07 (I2V) / $5.57 (editing board) | 720p | Yes | AA I2V leaderboard · AA editing leaderboard |
| Kling 3.0 Omni 1080p Pro | KlingAI | $16.80 | 1080p | Yes | AA I2V leaderboard |
| Kling 3.0 1080p Pro | KlingAI | $20.16 | 1080p | Yes | AA I2V leaderboard |
| Veo 3.1 | Google DeepMind | $24.00 | 1080p+ | Yes | AA I2V leaderboard |
| HappyHorse-1.0 | Alibaba-ATH | $27.04 | — (provider's default tier) | Yes | AA editing leaderboard |
| SkyReels V4 | Skywork AI | $37.50 | — (provider's default tier) | Yes | AA editing leaderboard |
| grok-imagine-video | xAI | $4.20 | — (provider's default tier) | Yes | AA I2V leaderboard |
Two things to notice:
- H3 is mid-tier on price, not the cheapest. Gemini Omni Flash ($6/min) and grok-imagine-video ($4.20/min) both come in under H3. If price-per-minute is your only metric, those are the cheaper picks — but neither currently matches H3's 12-reference-input cap or its #1 editing rank.
- H3 is far cheaper than the closed premium tier. Kling 3.0 Pro, Veo 3.1, HappyHorse-1.0, and SkyReels V4 are all 2x–4x the price of H3 while sitting below H3 on the editing board. For teams that compare against the premium closed names — not the budget segment — H3 is the price-performance winner on editing.
The pricing is volatile and re-checks monthly per our editorial freshness standard. Verify before quoting in a production budget.
How do you actually use MiniMax H3 for a video editing workflow?
The model is accessed via MiniMax's Platform API (model ID MiniMax-H3, asynchronous submit-then-poll), the consumer Hailuo AI app, or third-party inference hosts like fal.ai and Vercel's AI Gateway. The API supports four generation modes, but the editing and reference workflows are where H3 earns its leaderboard spot.
Step 1: Pick the right mode for the job
| Mode | Input | Use when |
|---|---|---|
| Text-to-Video | Prompt only | You're starting from zero with no visuals in hand |
| First/Last Frame | Prompt + 1 or 2 images | You know the starting and/or ending frame and want H3 to fill the motion between |
| Reference-to-Video (Omni Reference) | Prompt + up to 9 images, 3 videos, 3 audio | You need a consistent character, voice, style, or camera language across shots |
| Instruction-based Editing | Prompt + a finished clip | You have footage you like and want to change one element without redoing it |
Step 2: For editing, write the change as a targeted instruction
The mistake that wastes the model is treating editing like re-generation. Per the MiniMax docs, the editing prompt should describe what changes in the spot you point at — not rewrite the whole scene. Examples that work:
At 0:05, swap the coffee cup on the desk for a black laptop. Keep everything else identical.Relight the entire scene from bright afternoon to dusk with warm orange key light from camera left.Replace the line "Visit us today" at 0:09 with "Welcome to the team" in the same voice.Remove the parked car on the left edge of the frame. Fill the background with the same street tone.
The unedited parts of the frame stay pixel-stable, and the audio track stays in sync — including replacing a single line of dialogue while keeping the rest of the performance intact.
Step 3: For multi-reference work, use the full 12-file budget
Don't send one reference where you could send four — the consistency gains compound. A common reference stack for a brand intro clip:
- 1 image of your brand color palette + logo (controls visual style)
- 1 image of the character (face, wardrobe) — keeps identity on-model
- 1 short video of the camera movement you want the model to mimic (Hitchcock dolly, push-in, orbit)
- 1 audio clip of the voice you want cloned onto the character
Cite each in the prompt by order ("Image 1's brand colors, Image 2's subject walking, Video 1's camera arc, Audio 1's voice"). You can write up to 7,000 characters of prompt — long enough for a full shot list, not just a sentence. A vague two-line prompt produces a vague clip; the model earns its keep on long, specific prompts.
Step 4: Poll the task, then iterate on the result you already like
The API is asynchronous — submit, get a task_id, poll until succeeded, download from content.url. The cost advantage compounds when you edit the result you already like rather than regenerate from scratch each time. A working pattern:
- Generate the base clip with text-to-video or reference-to-video.
- Pick the take you like. Save it.
- Use instruction-based editing to spin variants: different on-screen text, different background, different closing line, relit version, re-voiced version.
- Each variant edits the base — no fresh generation cost, no new seed luck.
That's the loop H3 was optimized for. If your team churns through fresh generations because each iteration starts over, you're paying the model's worst case every time.
When is MiniMax H3 the wrong tool?
This is the part most coverage skips. The leaderboard says #1 in editing — and a lot of people read that as "#1 at AI video" and reach for H3 on every job. That's the mistake. H3 is overkill for a class of tasks a simpler, cheaper model handles better.
Pick a different model when:
- You only need one quick clip from a single line of text. Pure text-to-video with no reference inputs is the bulk of what Veo 3.1 Lite ($4.80/min), grok-imagine-video ($4.20/min), or an open model like LTX-2.3 Fast ($2.40/min) are for. H3's 12-reference slots and editing controls add setup time without adding value when you have nothing to reference.
- You need longer single-shot durations or a verified 4K option. H3 tops out at 2K and 15 seconds per generation. Longer sequences require multi-shot stitching or an extend feature. Veo 3.1 is the headline closed competitor that pushes longer single-shot durations; check current models' duration limits before committing.
- You need the absolute #1 T2V or I2V result. Artificial Analysis has Gemini Omni Flash ahead on T2V (1245 vs 1242) and both Seedance 2.0 and Gemini Omni Flash ahead on I2V. The Elo gaps are narrow, but on those exact tasks H3 is not the leader.
- You need an actually-released open model today. As of August 2, 2026, MiniMax has only stated the weights will be released "in the coming days, subject to applicable laws and regulations" — no weights, no license file, no model card is public yet. If you cannot ship without self-hosting, LTX-2.3 is the strongest open-weights video model available right now, even though it sits well behind H3 on the leaderboards.
To match the model to the job, answer two questions before touching any API: What state does my source footage need to start in? (a text description, an image, a clip, or a stack of references?) and What's the change in the next iteration? (whole clip from scratch, or one element of a clip I already like?). H3 wins the second answer most of the time and ties on the first.
Is MiniMax H3 actually open source?
Not yet, as of August 2, 2026 — but MiniMax has committed to opening the weights, and that commitment is the model's second differentiator after the editing leaderboard.
The official position, from MiniMax's launch post, is: "We plan to open up the model weights in the coming days, subject to applicable laws and regulations. Hardware compatibility has been a key consideration since the earliest stages of H3's design."
What's confirmed:
- The model architecture (
H3-Contextual Omni Representation,H3-VAE,H3-Omni Transformer,In-Context Regeneration) is described in the launch post. - Third-party hosting is already live on fal.ai and Vercel AI Gateway — but those are closed-API hosts, not weights.
- The MiniMax platform API and the consumer Hailuo AI app are live today.
What's not confirmed (treat as vendor claim until published):
- The license terms (reported as "MiniMax community license," free for organizations under $20M revenue per AlphaSignal — but the license file itself isn't published as of this writing).
- The weights repository (Hugging Face or ModelScope link).
- The parameter count, model card, or hardware requirements for self-hosting.
- Whether "in the coming days" is days, weeks, or longer.
If the weights land under the reported terms, H3 becomes the strongest open video model by a wide margin — the next-best open-weights models on Artificial Analysis (the LTX-2.3 family) sit roughly 200–260 Elo points behind it across the boards. That gap is large enough that an open-weight H3 would meaningfully change who can run production-grade video editing in-house. Until the weights land, "open" is a directional promise — use the API in the meantime.
How does MiniMax H3 fit the broader AI video landscape in 2026?
H3 launches into a competitive field that has been unusually busy in 2026. The Chinese labs are shipping faster than the leaderboards can settle: ByteDance's Dreamina Seedance 2.0 holds #1 on Image-to-Video (Elo 1196); Kuaishou's Kling 3.0 sits mid-board; MiniMax itself went public on the Hong Kong exchange in January 2026 per IndexBox, part of the group of well-funded 2022-founded startups labeled the "AI tigers".
| What changed in 2026 | Why it matters for H3 |
|---|---|
| Google launched Gemini Omni Flash (May 2026) | The model H3 trails on T2V by 3 Elo points — barely outside the confidence interval |
| ByteDance shipped Dreamina Seedance 2.0 (Mar 2026) | Holds #1 Image-to-Video on the AA leaderboard (Elo 1196); the strongest closed video generator on raw quality, beating H3 specifically on the I2V board |
| Kuaishou shipped Kling 3.0 Omni (Feb 2026) | Premium closed model; 2.5–3× H3's price-per-minute, sits mid-board |
| DeepSeek V4 Flash 0731 | Not a video model — but shows Chinese labs' open-weight + agentic pattern is now established across modalities |
| MiniMax's own IPO (Jan 2026, HKEX) | Public capital base per IndexBox; the open-weights promise becomes a credibility question on a listed company, not a startup |
The pattern: the closed premium names lead on quality; the budget names lead on price; H3 is the rare model that sits in the middle on both and leads only on the editing axis. For teams that need the editing axis more than the others, that's the right place to be.
What this means for you
You make ads, product videos, short social content, or brand intros — H3 is built for your workflow. Lean into the reference slots, write long prompts, and use instruction-based editing to spin ten variants of one winning clip instead of regenerating from scratch each time. That's where the price-performance lead actually shows up.
You need one quick clip from a single sentence — H3 is the wrong tool. Pick the cheaper, simpler model for that one task and revisit H3 only when your job starts needing consistency across shots, voice cloning, or iterating on existing footage. Powerful-tool-for-the-wrong-job is the failure mode the headline cycle won't warn you about.
You're a developer who can't ship without self-hosting — wait on H3's open-weights release. Use LTX-2.3 or the budget closed-API options in the meantime, and verify the actual license terms before committing engineering time — "open" in a tweet is not a license.
You're benchmarking AI video for production — read Artificial Analysis's leaderboards directly, not the headlines, and pay attention to the confidence intervals. The H3 "win" is editing-specific and within margin on T2V. Three Elo points is not a quality verdict; it is a tie.
For a broader walkthrough of the model's full feature set (T2V, I2V, first/last frame, the API step-by-step), see our complete MiniMax H3 guide. For tools and frameworks that pair well with H3 in a content-stack (the agent that calls the editing API, hallucination control on outputs), our AI agent operating system setup guide and the budget AI model decision guide cover the surrounding decisions.
FAQ
Q: Is MiniMax H3 the best AI video model? A: No — it's the best AI video editor, per Artificial Analysis where H3 ranks #1 in Video Editing with Audio (Elo 1130). It ranks #2 in Text-to-Video (behind Gemini Omni Flash) and #3 in Image-to-Video (behind Seedance 2.0 and Gemini Omni Flash). Read the editing win as #1 at editing, not #1 at video.
Q: How much does MiniMax H3 cost per video? A: Approximately $0.13 per second ($7.80/minute) for 2K output with native audio, per Artificial Analysis and MiniMax's blog. A 10-second 2K clip with audio runs roughly $1.30. Compare to $20.16/min for Kling 3.0 Pro or $24/min for Veo 3.1.
Q: Can MiniMax H3 edit a video I already have? A: Yes — instruction-based editing is its #1-ranked capability. You submit a finished clip and describe the change in natural language ("swap the sign at 0:05 to read 'OPEN'"), and only the region you named is edited while the rest stays pixel-stable. Audio stays in sync, including single-line dialogue replacement.
Q: How many reference inputs can MiniMax H3 take? A: Up to 12 files per request — 9 images, 3 video clips, and 3 audio clips (audio references must be paired with at least one image or video, per the platform docs). The 7,000-character prompt cap is large enough for a full shot list.
Q: Is MiniMax H3 free or open source? A: Not yet. The API is paid pay-per-use ($7.80/min at 2K). MiniMax has stated the model weights will be released under a "community" license "in the coming days" per the launch post, but as of August 2, 2026 no weights, license file, or hardware requirements have been published. The consumer Hailuo AI app offers free starting credits.
Q: Who should NOT use MiniMax H3? A: Anyone who only needs one quick clip from a single text prompt and has no reference material — a cheaper T2V model (grok-imagine-video at $4.20/min, LTX-2.3 Fast open at $2.40/min) is the better tool. H3 also caps at 2K and 15 s, so longer single-shot needs should look at Veo 3.1 or another model that pushes beyond 15 s before committing.
Q: Where can I access MiniMax H3? A: The MiniMax Platform API directly, the Hailuo AI consumer app, or third-party inference on fal.ai and Vercel AI Gateway. All four are live as of August 2, 2026.

Discussion
0 comments