Robots Atlas>ROBOTS ATLAS
Artificial Intelligence

MiniMax H3: an open video model aiming to be multimedia's DeepSeek

Sir Robot4 August 2026 · 3 min read
MiniMax H3: an open video model aiming to be multimedia's DeepSeek

MiniMax has released H3, a multimodal generation model that unifies the understanding and creation of text, images, video and audio in a single system. Announced on July 31, 2026 on the company blog, H3 is positioned to differ from closed rivals such as Sora, Seedance and Keling through a promised open-weights release.

Key takeaways

  • H3 generates video up to 15 seconds at 2K resolution with native stereo audio produced jointly with the image
  • Per Artificial Analysis rankings (as of July 31, 2026), H3 placed first in video editing, second in text-to-video and third in image-to-video
  • 2K output is priced at ¥0.8 (about $0.11) per second — under one-third of competitors' rates
  • A proprietary H3-VAE encoder cuts sequence length roughly 4× versus rivals
  • Model weights are set for open download "in the coming days"

One model instead of separate tools

Most video generators do one thing: turn text into a clip, or an image into motion. H3 goes wider — it processes text, image, video and audio in one context and returns finished audio-video output, rather than bolting sound on at the end. In practice that means transferring motion from one clip to another (V2V: transferring motion from one video clip to another — video-to-video) or keeping subtitles, logos and on-screen text intact despite changes in perspective, lighting and camera movement — a weak spot for most generative models.

The company describes H3 as "DeepSeek for multimedia." The reference points to a strategy where the edge comes not from a single best score but from combining openness with low cost. The same logic drove Chinese language models that pushed inference costs to levels closed players were unwilling to match.

Scores and pricing

H3's strongest case is video editing — Artificial Analysis puts the model in the global lead here, ahead of closed alternatives. In the other categories it sits near the top, placing it in the highest tier without dominating each one.

CategoryH3 rank
Video editing1st
Text-to-video2nd
Image-to-video3rd

Price matters here as much as quality. 2K output costs under one-third of what competitors charge, and 768p material comes in below half the typical 720p rate. Behind that economy sits the technical layer: the H3-VAE encoder shortens the sequence roughly fourfold, a Contextual Omni Representation mechanism cuts context from about 100K tokens to roughly 4K on average, and the H3-Omni Transformer architecture separates understanding and generation workloads, lifting training throughput by about 30%.

¥0,8 / s2K output price — under one-third of rivals' ratesMiniMax

The market reaction was immediate — the company's Hong Kong-listed shares jumped over 10%, to HK$249.4, as reported by tmtpost.

Why it matters

Video generation has so far been the domain of closed, expensive systems. If MiniMax genuinely opens the weights of a flagship video model, it shifts the reference point for the whole category — much as open language models forced the market into price cuts. The combination is what counts: competitive quality in video editing, pricing under one-third of rivals' rates, and the option to run locally. For ad studios, e-commerce and game creators, that means potential control over cost and data that a closed API cannot offer. Openness, however, remains a promise, not yet a fact.

What's next?

  • MiniMax said it would release H3's weights "in the coming days" from July 31, subject to applicable regulations — the actual timing and license will decide adoption scale
  • An open-weights release would let outsiders independently verify the claimed Artificial Analysis scores and inference costs on their own hardware

Sources

Share this article