
Developers & API integrators
Integrate via MiniMax API (`MiniMax-H3`), Vercel AI Gateway, fal.ai, or self-host open weights. Single architecture handles T2V, I2V, reference, and editing.
MiniMax H3 is a general-purpose omni-modal generative system — a 33B-parameter model that understands text, images, video, and audio in unified context, generating 2K video with native stereo sound in a single forward pass.
Generate a video to see the preview in this area
MiniMax H3 breaks closed-source dominance with open weights, industry-leading price-performance, and a task-generalization architecture built for the open ecosystem.
H3-Base weights released on Hugging Face under the MiniMax Community License — permitting commercial use for organizations under $20M revenue. Build customized versions on diverse AI hardware.
A 33B-parameter dense single-stream Transformer with ~20B active inference parameters. Uses Qwen3-VL-32B as encoder, 3D multimodal RoPE, and separates understanding/generation compute for 30% training throughput gain.
Language acts as the bridge unifying tasks into open, descriptive form. Describe relationships between multimodal context and target video in natural language — H3 handles full-modality understanding.
Instead of conventional super-resolution, H3 regenerates its own 768p output in-context at 2K — recovering fine details like small text that upscalers can't restore.
Follow these steps to start generating with MiniMax H3 using the generator on this page or the MiniMax API.
Select text-to-video for prompt-only generation, first/last-frame for image-to-video, or reference-to-video for omni-reference workflows with up to 12 mixed input files.

MiniMax H3 serves developers, studios, and creators who need open, controllable, production-grade AI video with multimodal understanding.

Integrate via MiniMax API (`MiniMax-H3`), Vercel AI Gateway, fal.ai, or self-host open weights. Single architecture handles T2V, I2V, reference, and editing.

Excel at instruction following, accurate text and brand rendering, and V2V motion transfer. Built for ads, product design, UI/UX, and gaming content.

Fine-tune, customize, and deploy on your own hardware. Compatible with a broad range of AI accelerators — hardware compatibility was a key design consideration.
Write a prompt or attach multimodal references, and MiniMax H3 generates 2K video with native stereo audio — dialogue, SFX, and ambience in one pass.
MiniMax H3's task-generalization design delivers broad multimodal capabilities from the pretraining stage.
One transformer handles text-to-video, image-to-video, reference-to-video, video editing, and audio generation — no separate T2V, I2V, and TTS models stitched together.
Joint video and audio generation in one forward pass at 32kHz stereo. Voice, sound effects, and music jointly modeled — no separation between audio domains.
Output durations from 4 to 15 seconds at 24 frames per second. Aspect ratios from 2:5 to 5:2, covering widescreen cinematic to vertical social formats.
Completely overhauled tokenizer with across-the-board gains in reconstruction quality. 4x effective sequence length gain — key technology behind native 2K support.
Supports Arabic, Chinese, English, French, German, Italian, Japanese, Korean, Portuguese, Russian, and more — with varying degrees of audio language support.
At 2K, per-second price is less than a third of mainstream models (~0.80 CNY/sec). At 768p, less than half the price of mainstream 720p offerings.
Ranked as the world's most powerful AI model in video editing on Artificial Analysis benchmarks — trailing only in pure T2V against Gemini and Seedance.
From MiniMax H3's open omni-modal architecture to Wan 3.0's 30-second generation and Google Veo's cinematic quality.

Open-weight 33B omni-modal model with native 2K stereo audio and instruction-based editing.

MiniMax's consumer-facing brand for H3 — native 2K, synced audio, and omni-reference control.

Alibaba's 30-second multimodal video model with document-to-video conversion.

Google DeepMind cinematic AI video with 4K and native audio.
Deep dives into H3's architecture, API integration, and creative workflows.
Still have questions? We're here to help.