MiniMax has shipped an AI video model that no longer treats sound as an afterthought. On July 31, 2026, the Shanghai-based company released MiniMax H3, a general-purpose model that ingests text, images, video, and audio together and returns clips of up to 15 seconds at native 2K resolution (2,560x1,440 pixels, 24fps) with synchronized stereo sound generated in the same inference pass as the picture.
MiniMax describes H3 not as a text-to-video tool with bolted-on extras but as an omni-modal system built from a single pretraining paradigm. "H3 understands unified context across text, images, video, and audio, generating video with native stereo sound, up to 15 seconds at 2K resolution," the company said in its launch post, adding that reference and editing instructions -- swap a product, match a character's voice to a supplied audio clip, copy a camera move from a reference video -- are now expressed in natural language rather than routed to separate expert models for each task.
Reuters was first to report the release outside MiniMax's own channels. "Chinese AI firm MiniMax released a new video-generation model on Friday that can process text, images, video and audio, stepping up competition in a fast-growing market led by rivals ByteDance and Kuaishou," Reuters' Eduardo Baptista wrote from Beijing, noting that MiniMax also plans to publish H3's model weights "within days," extending the open-weight approach it has used for its M-series language models into video generation for the first time among major Chinese labs.
The economics are the headline for developers. MiniMax says 2K generation costs less than a third of what mainstream rivals charge, and independent tracking firm Artificial Analysis pegged the API rate at $0.13 per second -- about $7.80 per minute -- versus roughly $22.45 per minute for ByteDance's Seedance 2.0 at 1080p and $20.16 per minute for Kuaishou's Kling 3.0. Google's Gemini Omni Flash still undercuts H3 slightly at $6.00 per minute and, per Artificial Analysis, currently outranks H3 on both text-to-video and image-to-video preference scores; H3's clear win is video editing, where it holds the top Elo score (1,130, drawn from more than 5,000 blind-preference comparisons) on the firm's leaderboard. The model accepts up to nine reference images, three reference videos, and three reference audio clips per request -- fewer than the 50 combined inputs ByteDance's Seedance 2.5 offered when it launched the very same day, underscoring how compressed the release calendar among Chinese video labs has become.
The technical unlock behind the price cut is a rebuilt tokenizer, H3-VAE, which MiniMax says delivers a four-times gain in effective sequence length over its prior architecture, cutting the compute cost of high-resolution frames enough to make 2K the default rather than a premium tier. Paired with what MiniMax calls In-Context Regeneration -- a technique that lets the model re-read its own multimodal context to sharpen a low-resolution draft into full 2K output instead of relying on a bolted-on super-resolution pass -- the company says small text, product logos, and facial detail survive the upscale far better than with conventional pipelines. A new H3-Omni Transformer, which MiniMax says lifted end-to-end training throughput by nearly 30%, replaces the architecture used in the prior Hailuo-02 generation specifically to handle the wider variance in sequence length that multimodal inputs introduce.
Why It Matters
H3's launch lands in the middle of the most crowded stretch the AI video market has seen: ByteDance's Seedance 2.5 arrived the same day, Kuaishou's Kling 3.0 shipped earlier this year, and OpenAI pulled Sora from general availability in April after a short run. That leaves Chinese labs -- MiniMax, ByteDance, and Kuaishou -- setting the pace on price and multimodal capability while Google's Gemini Omni Flash is the main non-Chinese model still near the top of the leaderboard. MiniMax's promise to open-source H3's weights, even under a license that restricts free commercial use to companies with under $20 million in annual revenue, would make it the most capable open-weight video model released to date, a notable break from an industry where the leading systems have stayed proprietary. It also arrives as MiniMax, which raised roughly $619 million in a Hong Kong IPO at about a $6.5 billion valuation in January 2026, works to differentiate a commercial video business that sits on the same Hailuo platform currently facing an active US copyright suit from Disney, Universal, and Warner Bros. Discovery -- a case that survived a motion to dismiss in May and remains in discovery.
What to Watch
Watch whether MiniMax actually ships H3's weights on the promised timeline, and under what license terms larger enterprises will be asked to pay for commercial use. Watch Artificial Analysis's leaderboard for H3's text-to-video and image-to-video Elo scores to stabilize as more blind-preference votes accumulate -- today's second- and third-place rankings in those categories are still provisional. And watch how the Disney-Universal-Warner Bros. lawsuit against MiniMax's Hailuo platform, along with the pending Andersen v. Stability AI trial in September, shapes enterprise willingness to build production pipelines on Chinese video models regardless of price.
"Chinese AI firm MiniMax released a new video-generation model on Friday that can process text, images, video and audio, stepping up competition in a fast-growing market led by rivals ByteDance and Kuaishou."- Eduardo Baptista, Reuters, Beijing