The Tech ArchiveThe Tech ArchiveThe Tech Archive
Small BusinessMarketingDevelopers
ArticlesTopicsSeriesAbout

Get the practical AI brief

Verified, no-hype AI tips you can actually use - in your inbox. Free.

No spam. We verify what we send. Unsubscribe anytime.

The Tech ArchiveThe Tech Archive

The Tech Archive

AI news, analysis & explainers

AboutSmall BusinessMarketingDevelopersArticlesTopicsSeriesMethodologyAI DisclosureCorrections

© 2026 All rights reserved.

Back to home
0 readers reading
  1. Home
  2. Articles
  3. Artificial Intelligence
  4. MiniMax H3 (2026): The Open-Weights AI Video Model That Generates 2K Video With Audio at One-Third the Price

Contents

MiniMax H3 (2026): The Open-Weights AI Video Model That Generates 2K Video With Audio at One-Third the Price
Artificial Intelligence

MiniMax H3 (2026): The Open-Weights AI Video Model That Generates 2K Video With Audio at One-Third the Price

MiniMax H3 is an omni-modal AI video model that produces 2K clips with native stereo audio, edits existing footage on instruction, and costs $0.13/sec — about a third of mainstream rivals.

Sham

Sham

AI Engineer & Founder, The Tech Archive

19 min read
1 views
August 4, 2026

Verdict: MiniMax H3 is the first frontier-class AI video model to promise open weights, generate 2K clips with native stereo audio in the same pass, and edit existing footage from a natural-language instruction — all at roughly one-third the per-second price of mainstream rivals like Sora 2 Pro. Released on July 31, 2026, by Shanghai-based MiniMax (the company behind the Hailuo video brand), it already ranks #1 on Artificial Analysis's independent Video Editing leaderboard and top-three on its Text-to-Video and Image-to-Video boards.

For advertisers, e-commerce teams, and small studios that live or die by short, sound-ready clips, H3 collapses what used to be three or four separate tools — a generator, an audio model, a super-resolution module, and a video editor — into one API call. The catch: at the time of writing, the open weights are promised "in the coming days" but not yet on Hugging Face, and the license is a MiniMax Community License (commercial use only for organizations under $20M in revenue), not an OSI-approved open-source license. Treat the current API version as a product; treat the open-weights release as an expectation.

TL;DR

  • What it is: MiniMax H3, an omni-modal video model that takes text, ≤9 images, ≤3 videos, and ≤3 audio clips as context and returns 2K video with native stereo sound.
  • Best for: Teams that need commercial-ready short clips with sound, in-place editing, and reference-driven control — not raw 4K fidelity.
  • Price: $0.13 per generated second at 2K (≈$1.95 per 15-second clip); 768p tier at $0.09/sec; audio input free; first 5 reference images free.
  • Leaderboard: #1 Artificial Analysis Video Editing (Elo 1,130), top-3 Text-to-Video and Image-to-Video (independently measured, blind votes).
  • Open weights: Announced, not yet shipped. MiniMax Community License — free for non-commercial and commercial use under $20M revenue, attribution required.

Last verified: 2026-08-04. Pricing, limits, weight availability, and leaderboard rank change fast — re-check before budgeting.


What makes MiniMax H3 different from every other AI video model?

H3 is not "another text-to-video generator." It is an omni-modal model: a single neural network that understands text, images, video, and audio in one shared context window, and that can both generate new clips and edit existing ones from a sentence.

Most existing video systems farm each task out to a separate specialist — a text-to-video model, a separate image-to-video model, a separate audio model, a separate super-resolution module, and yet another editor for swaps and motion transfer. H3 puts all four "languages" in one head, which is exactly why you can say "use the camera move from Video 1, have the character in Image 2 sing the vocals in Audio 3" and the model resolves the cross-modal relationships itself (MiniMax, H3 blog post, 2026-07-31). The official architecture diagram on Hugging Face confirms it: a single H3-Omni-Transformer jointly predicts video and audio latents from a packed multimodal sequence, rather than stitching outputs from separate encoder-decoder pairs (HuggingFace MiniMaxAI/MiniMax-H3 model card).

The practical payoff is that you stop passing notes between four translators. A prompt that used to mean "generate in Tool A, dub audio in Tool B, upscale in Tool C, and re-cut in Tool D" now resolves to one API call — and the model's editing capability means most revisions don't require regenerating the whole clip from scratch.

What can MiniMax H3 actually do? (In-place editing is the headline)

The H3 capability set splits into three buckets: generation, reference control, and editing. The editing half is what puts H3 at #1 on the Artificial Analysis Editorial leaderboard — most competitors do not seriously compete in instruction-based editing at all.

  • Text-to-video with native stereo audio (32 kHz) generated in the same pass, not dubbed on later.
  • Image-to-video with up to 9 reference images for locked identity, style, or product consistency.
  • First-and-last-frame transition: supply opening and closing images, and H3 generates the in-between motion.
  • Video editing by instruction: "make it nighttime," "remove the person on the left," "replace the product with the one in Image 3" — only the requested change updates; the rest of the shot is preserved.
  • Motion transfer: take camera pans, blocking, and timing from an existing clip and apply them to a different subject in a new clip.
  • Voice and audio references: pipe in a reference audio clip for voice timbre, pacing, or synced lip movement.

Verified output specs from the official MiniMax platform documentation:

Spec Value
Native resolution 2K (768P tier also priced)
Clip length 4–15 seconds, integer-second durations
Frame rate 24 FPS
Audio 32 kHz native stereo, same-pass
Omni-reference inputs ≤9 images, ≤3 videos (2–15 s each), ≤3 audio (must accompany an image/video); max 12 files, 7,000-character prompt
Supported languages (audio) 11 stable: Arabic, Chinese, English, French, German, Italian, Japanese, Korean, Portuguese, Russian, Spanish
Aspect ratios 21:9, 16:9, 4:3, 1:1, 3:4, 9:16, plus adaptive

H3's durability sweet spot is exactly 5–15 seconds — the length of a TikTok ad, an Instagram Reel product shot, a YouTube bumper, or a hero video on a landing page. For longer narrative work you'd stitch multiple calls, but the single-call use case is where most commercial short-form video already lives.

How does H3 compare to Sora 2 Pro, Veo 3.1, Kling 3.0, and Seedance 2.0?

H3 does not lead on raw resolution. Kling 3.0, Veo 3.1, and Seedance 2.0/2.5 already offer native 4K (Kling Ultra and Veo Pro) or 2K (Seedance Pro). What H3 leads on is the combination of 2K + same-pass audio + heavy reference control + a price roughly a quarter to a third of the closest Western alternative. The comparison below aggregates verified per-second prices from each vendor's published pricing docs and third-party aggregator rates where no official rate exists (AIReiter, "MiniMax H3: Specs, Pricing, and How It Compares," 2026-07-31); (Artificial Analysis Video Editing Leaderboard).

Model Max resolution Max single clip Native audio Reference inputs API $/sec (verified) Editing (instruction-based)
MiniMax H3 2K 15 s Yes, same-pass stereo 9 img + 3 vid + 3 aud $0.13 (2K), $0.09 (768p) #1 on Artificial Analysis
Google Veo 3.1 up to 4K ~8 s Yes, 48 kHz dialogue Limited from $0.03 (Lite, no audio) Limited
Kuaishou Kling 3.0 native 4K ~10 s Yes, lip-sync 5 languages Limited (Elements 3.0) "Varies" — no published per-sec rate Limited
ByteDance Seedance 2.0 / 2.5 up to 2K 10 s Yes, multilingual 9 img + 3 vid + 3 aud (omni-ref) "Varies" — no published per-sec rate Mid-table (Elo 1,035, 720p)
OpenAI Sora 2 Pro 1080P up to ~25 s Yes Limited ~$0.50 (third-party aggregator rate) Not a focus area
Hailuo 2.3 (predecessor) 1080P ~10 s No native audio None ~$0.082 N/A

The take-away: if your bottleneck is "I need 10 good 15-second product clips, with sound, by Friday, and I don't want to spend $500," H3 is currently the only published-rate model that hits that bar. Veo 3.1 Lite is cheaper per second but omits audio entirely, so it's a different product. Kling 3.0 leads on 4K fidelity, Seedance 2.0 on director-style multi-shot narrative; neither has published a comparable per-second API rate at the time of writing.

Worth tracking the broader China-AI pattern behind H3: MiniMax's release ships in the same window as dozens of open-weights model launches piling pricing pressure on closed labs. We've tracked how China's open-weight AI models are forcing Anthropic and OpenAI to compete on price elsewhere on the blog — H3 is the same pattern in video form.

The four technologies behind H3 (in plain language)

MiniMax's launch materials name four architectural pieces. The names sound academic; the function is concrete.

  1. H3-Contextual Omni Representation. A captioning pipeline that turns raw, unstructured multimodal input into a single descriptive prompt the model can act on. The pipeline does roughly 100K tokens of inference on the source material and distills it down to an average ~4K tokens — MiniMax describes language as the "generalizable bridge" that lets H3 handle editing relationships it was never explicitly trained on (MiniMax H3 blog). This module (called H3-Context-IR) is hosted and not part of the open-weights release; MiniMax recommends you either use their API or follow their Prompting Guidance to build your own (HuggingFace model card).

  2. H3-VAE. A complete overhaul of the visual tokenizer, with a 16× spatial / 4× temporal compression ratio and 24 latent channels. MiniMax claims a 4× gain in effective sequence length vs. the prior Hailuo 02 tokenizer — this is the technical reason native 2K is affordable at all. Without H3-VAE, the token budget for a 2K clip would be prohibitive.

  3. H3-Omni Transformer. The core 33B-parameter dense transformer. It's "omni" because attention and FFN layers contain no modality-specific structure — modality specialization lives only in the input encoders, output decoders, and AdaLN (adaptive layer-norm) branches. The AdaLN outputs can be precomputed and cached, so they don't need to be loaded for inference-only deployments. MiniMax reports a ~30% lift in end-to-end training throughput from splitting understanding and generation workloads (MiniMax H3 blog).

  4. H3-In-Context Regeneration. Instead of bolting on a separate super-resolution model after the low-res generation, H3 regenerates its own low-res output in-context — re-reading the original multimodal references to recover fine details, small text, and brand elements. The advantage over conventional upscaling: the model can hallucinate less because it's grounding against the original context, not inventing detail it never saw. The H3-Regenerate-2K module is not yet open-sourced either (HuggingFace model card).

What's actually being released as open weights is H3-Base, in two variants: FL2VA (first-and-last-frame) and Ref2VA (omni-reference). The context-processing and 2K-regeneration layers stay as a hosted service for now. That's an important caveat for anyone planning to self-host and expects to drop in a 2K-capable model: you'd be running the base 768p output unless you wire in the hosted or your own context-orchestration layer.

For similar context on why open-weights releases are reshaping how builders pick models, see our breakdown of Qwen 3.8 Max: what Alibaba's 2.4T open-weight model actually delivers.

How much does MiniMax H3 cost? (verified per-second rates)

MiniMax publishes per-second rates directly on its pricing page. These are the verified numbers as of July 31, 2026:

  • 2K resolution: $0.13 per generated second → a 15-second clip costs ≈$1.95.
  • 768P resolution: $0.09 per second.
  • Audio input: free.
  • Reference images: first 5 free; each additional image (up to 9 total) is $0.04.

MiniMax's own claim, restated on the launch blog: "At 2K, H3's per-second price is less than one-third of mainstream models, and at 768p, it's less than half the price of mainstream models' 720p" (MiniMax H3 blog). The numbers from third-party aggregators bear that out: Sora 2 Pro at 1080P is reported at approximately $0.50/second via aggregator pricing, so H3's 2K rate at $0.13/s is about a quarter of that figure (AIReiter pricing analysis). Veo 3.1 Lite starts at around $0.03/second but omits native audio, so it's a different product category.

The honest read on cost: H3 is the cheapest model that ships both 2K resolution and native stereo audio in the same pass. The one cheaper credible rival (Veo Lite) is not a comparable product because it has no audio; everything with comparable audio and resolution is more expensive.

Pricing is volatile. Treat every number above as checked-on-August-4-2026 and re-verify before any procurement or budget commit — see the Updates & Corrections log at the bottom for our own re-check cadence.

Is MiniMax H3 open source? (the important nuance)

Not yet, and "open source" is the wrong label even when it lands. MiniMax has committed in writing to publishing H3's weights "in the coming days, subject to applicable laws and regulations" (MiniMax H3 blog) — Reuters independently confirmed the timeline (NYU Shanghai RITS, 2026-07-31). As of August 4, 2026, no H3 video repository exists on the MiniMaxAI Hugging Face organization; what's there today is a model card and documentation, not downloadable weights.

When the weights ship, they will land under the MiniMax Community License: free for non-commercial use; commercial use permitted only for organizations with annual revenue under USD $20M; attribution required. That is a source-available commercial license with a revenue ceiling, not an OSI-approved open-source license — the same structural shape Meta used for Llama and that Krea 2 uses. For an indie studio or a Series A startup, it's functionally free; for a $50M/yr ad production company, it's not, and you should verify rights before integration (SaaSCity analysis, 2026-07-31); (DreamPixelForge analysis).

Two further asterisks on the "open" claim:

  • H3-Context-IR (the multimodal context-orchestration system MiniMax says is "critical" to output quality) is hosted-only and not in the open-source release (HuggingFace model card). You can use the API or build your own context processor; you cannot self-host theirs.
  • H3-Regenerate-2K (the in-context module that produces the 2K tier) is also not yet open-sourced. Self-hosted H3-Base runs at 768p by default; reaching 2K without the hosted module is an open research problem.

The open-weights release as currently scoped gets you a 33B-parameter base model that does text-to-video and reference-to-video at 768p, with full attention (sparse attention is "in a future update"). For full Hailuo-grade 2K-with-audio generation, you still need the API.

What are the Artificial Analysis leaderboard numbers, and should you trust them?

H3 ranks #1 on the Video Editing (With Audio) leaderboard with an Elo of 1,130 from over 5,000 blind human-preference votes, narrowly ahead of Google's Gemini Omni Flash (Elo 1,122) and Alibaba's HappyHorse-1.0 (Elo 1,096). On Text-to-Video (With Audio), it ranks #2 (behind Gemini Omni Flash). On Image-to-Video, top-three (Artificial Analysis Video Editing Leaderboard, accessed 2026-07-31); (ExplainX, 2026-07-31).

These numbers are independently measured — Artificial Analysis uses blind human-preference voting, the same methodology LMArena uses for text models — not vendor-reported benchmarks. That matters. We've written about how AI benchmark gaming works and why leaderboard scores can lie, and the Artificial Analysis methodology is one of the few we defend against the usual "vendors cherry-pick their own evals" critique. The sample size of 5,000+ blind votes at launch is meaningful.

Caveats, in the interest of honesty:

  • The rankings are a launch-day snapshot and will move as more votes accumulate.
  • Editing is where H3 genuinely leads; on raw single-shot generation, the field is closely contested and H3 does not claim a clear #1.
  • The Elo differential between H3 and Gemini Omni Flash on editing is 8 points — within the typical noise band for these leaderboards. Treat it as "tied for the lead," not "decisively best."
  • H3 is one of the only models on the editing leaderboard that accepts a source video as input at all; competitors that don't compete on editing aren't strictly comparable.

What does this mean for you?

If you run ads, e-commerce, or product video at scale: H3 is worth a serious trial immediately, because the per-second economics genuinely change the iteration math. At ~$1.95 per 15-second 2K clip you can afford 50 generations to find 5 great ones; at the Sora 2 Pro aggregator rate that same search costs roughly $375. Iteration cost is itself a quality mechanism — cheap retries surface better clips.

If you're an indie studio under $20M revenue: Treat the upcoming weights as a "worth planning around" signal but not a current deployment target. The current API is what you can build on; the open release, when it lands, unblocks private/custom-fine-tuned self-hosting at 768p. For automated video pipelines embedded in agents, you can already drive MiniMax (or a competitor like Higgsfield) through MCP from an agentic workflow — our Higgsfield MCP guide for AI agent video and image generation shows the pattern.

If you're an enterprise over $20M revenue: The MiniMax Community License caps mean you'll need to negotiate a commercial license with MiniMax before integrating H3 into a product. The model is technically impressive but the rights question is real; verify terms directly with MiniMax before building on the weights.

If you're tracking AI-generated content provenance: Note that H3-generated clips fall under the same disclosure expectations as any other AI-generated media. Under the EU AI Act Article 50 transparency rules (in force since August 2026) businesses must disclose when content is AI-generated where it could be mistaken for authentic human work — see our Article 50 compliance guide for the scope.

If you're benchmarking or writing about H3: Treat vendor demo reels as marketing, not measurement. The only apples-to-apples comparisons are prompt-matched runs you generate yourself or independent leaderboards with sufficient sample size. The Editing leaderboard number is the one claim worth citing; "H3 beats Sora 2" headlines are overclaims.

MiniMax H3 FAQ

Q: What is MiniMax H3? A: MiniMax H3 is an omni-modal AI video model released by Shanghai-based MiniMax on July 31, 2026. It takes text, images, video, and audio as context in a single request and returns 2K clips up to 15 seconds long with native stereo audio, and can edit existing footage from a natural-language instruction.

Q: How much does MiniMax H3 cost? A: $0.13 per generated second at 2K resolution (≈$1.95 for a 15-second clip) and $0.09 per second at 768p, via the official MiniMax platform API. Audio input is free; the first 5 reference images are free, with each additional image up to 9 total costing $0.04.

Q: Is MiniMax H3 open source? A: Not yet, and "open source" is the wrong term even when it ships. MiniMax has committed to releasing the H3-Base weights under the MiniMax Community License — free for non-commercial use, commercial use allowed only for organizations under $20M annual revenue, attribution required. That is a source-available license with a revenue cap, not an OSI-approved open-source license. As of August 4, 2026, the weights are not yet downloadable.

Q: How does H3 compare to Sora 2 Pro, Veo 3.1, and Kling 3.0? A: H3 is the first frontier video model to combine 2K + native stereo audio + heavy reference control + published per-second pricing under any of those rivals (about a quarter of Sora 2 Pro's rate and a third of "mainstream" models per MiniMax's claim). Veo 3.1 and Kling 3.0 lead on raw resolution (4K); Seedance 2.0 leads on multi-shot narrative; H3 leads on independent blind-vote rankings for instruction-based video editing. None of Kling, Seedance, or Veo has published a comparable per-second API rate at the time of writing.

Q: Can I edit existing videos with MiniMax H3? A: Yes — instruction-based editing of existing footage is H3's strongest differentiator. You can say "replace the product in the actor's hand with the one in Image 3," change lighting to night, remove a person from the shot, or transfer camera motion from one clip to a different subject — only the requested change updates. This is the capability that puts H3 at #1 on the Artificial Analysis Video Editing leaderboard.

Q: What are MiniMax H3's maximum clip length and resolution? A: 15 seconds per clip at 2K resolution (with a cheaper 768p tier), 24 frames per second, native 32 kHz stereo audio, integer-second durations from 4 to 15 seconds. Aspect ratios supported: 21:9, 16:9, 4:3, 1:1, 3:4, and 9:16, plus adaptive. Reference inputs: up to 9 images, 3 videos (each 2–15 s), and 3 audio clips per generation, with a 12-file mixed maximum.

Q: When will MiniMax H3 weights be released? A: MiniMax said "in the coming days" as of the July 31, 2026 launch; Reuters independently confirmed the same timeline. What will be released is H3-Base (FL2VA and Ref2VA variants only); the H3-Context-IR context-orchestration layer and the H3-Regenerate-2K module will remain hosted-only. As of August 4, 2026, no weights are on Hugging Face. Treat "open weights" as a stated intention, not a shipped artifact, until files appear.

Sources
  • MiniMax Research. "MiniMax H3: An Open Model Breaking the Boundaries Between Tasks and Modalities." July 31, 2026.
  • Hugging Face. "MiniMaxAI/MiniMax-H3 model card." Accessed August 4, 2026.
  • MiniMax Platform Documentation. "Video Generation guide." Accessed August 4, 2026.
  • Artificial Analysis. "Video Editing Leaderboard (With Audio)" and "Text-to-Video Leaderboard." Accessed August 4, 2026.
  • AIReiter. "MiniMax H3: Specs, Pricing, and How It Compares." July 31, 2026.
  • SaaSCity. "MiniMax H3: The Open-Weights Video Model That Just Undercut Everyone by 3x." July 31, 2026.
  • NYU Shanghai RITS Library. "MiniMax Releases H3: 2K Video With Native Audio, Open Weights Promised." July 31, 2026.
  • DreamPixelForge. "MiniMax H3 (Hailuo 3.0): 2K Video, Native Audio, and Open Weights." August 1, 2026.
  • ExplainX. "MiniMax-H3: An Open Video Model Ranked #1 on Video Editing With Audio." July 31, 2026.
Updates & Corrections
  • 2026-08-04 — First published. All specs, prices, leaderboard ranks, and license terms verified against primary sources as of this date. Open-weights release marked as announced-but-not-shipped.
  • Re-verification cadence (per our editorial standard): Pricing/limits/versions will be re-checked monthly. Open-weights availability will be re-checked weekly until the files appear on Hugging Face. Leaderboard rankings re-check quarterly or on a release event from a competing lab.

Get the practical AI brief

Verified, no-hype AI tips you can actually use - in your inbox. Free.

No spam. We verify what we send. Unsubscribe anytime.

Tags

#minimax-h3#text-to-video#hailuo#ai-video-generation#open-weights-model#multimodal-ai

Discussion

0 comments
Sham

Sham

AI Engineer & Founder, The Tech Archive

AI engineer (Azure AI-102/AI-900). Writes practical, tested, hype-free guides on using AI for real work and small business at The Tech Archive.

Related Articles

View all
Hermes Computer Use: How to Set Up a Background AI Agent That Works While You Keep Typing (2026)
Artificial Intelligence

Hermes Computer Use: How to Set Up a Background AI Agent That Works While You Keep Typing (2026)

14 min
EU AI Act Article 50: What Every Business Must Do About Chatbot and AI Content Disclosure (August 2026)
Artificial Intelligence

EU AI Act Article 50: What Every Business Must Do About Chatbot and AI Content Disclosure (August 2026)

15 min
Qwen 3.8 Max (2026): The 2.4T Open-Weights Model That Competes With Claude Opus 5 — Specs, Prices, and How to Actually Use It
Artificial Intelligence

Qwen 3.8 Max (2026): The 2.4T Open-Weights Model That Competes With Claude Opus 5 — Specs, Prices, and How to Actually Use It

13 min
How to Run a Fleet of AI Coding Agents for Free With Orca in 2026 (Parallel Worktrees, Compared)
Artificial Intelligence

How to Run a Fleet of AI Coding Agents for Free With Orca in 2026 (Parallel Worktrees, Compared)

16 min
How to Build an AI Job Search Agent With Claude Code in 2026 (The 29K-Star Open-Source Framework, Explained)
Artificial Intelligence

How to Build an AI Job Search Agent With Claude Code in 2026 (The 29K-Star Open-Source Framework, Explained)

18 min
Inkling-Small (2026): The Open-Weights Model That Matches Its 975B Sibling at a Quarter the Size
Artificial Intelligence

Inkling-Small (2026): The Open-Weights Model That Matches Its 975B Sibling at a Quarter the Size

11 min