MiniMax H3: The Open-Weight Multimodal Video Model, Explained

MiniMax H3 takes text, images, video clips, and audio into one shared context and returns 4-15 second clips at 24 FPS with native stereo sound - generation, referencing, and editing in a single model. This guide collects its real specifications, the documented prompt-writing rules, and every legitimate way to run it, then lets you try the same prompt style on AI Fruit's own video models in the generator below.

MiniMax H3: The Open-Weight Multimodal Video Model, Explained
Generate with AI Fruit's video models

从模板开始

选一个创意,再改成你的版本。

上传水果视频的参考图片

控制视频最终画面的呈现。

尽量说清动作、构图、光线和氛围。612 / 2000
s
登录后查看积分获取更多积分
Your Fruit Video
已恢复

准备生成你的第一条视频

选择模板或填写提示词,生成后可在这里预览和下载视频。

MiniMax H36768P16:932 积分
4-15sClip Length
768p / 2KResolution
StereoNative Audio
Open WeightsLicense

Ways to Run MiniMax H3

H3 is available through MiniMax's own app and API, through hosted providers, and as downloadable weights. Pick the route that matches how much control you need.

Official app and API

MiniMax ships H3 as a web app at hailuoai.video and a developer API at platform.minimax.io. This route gives you the full official pipeline, including the hosted context processing that turns free-form prompts and references into generator-ready instructions.

  • hailuoai.video web app
  • platform.minimax.io API
  • Full official pipeline

Hosted on fal

fal hosts H3 as a launch partner across three endpoints - Text to Video, Image to Video, and Reference to Video - published under the hailuo-03 API name. fal lists $0.26 per second of 2K output with no subscription, so short tests stay cheap.

  • Three hosted endpoints
  • $0.26/s at 2K on fal
  • No subscription needed

Open weights on Hugging Face

The weights live at MiniMaxAI/MiniMax-H3 under the MiniMax H3 Community License, with a Diffusers pipeline for local inference. The open release covers the base generator and the 2K regeneration pass; the hosted context-processing system is not part of it, so local runs follow the published prompting guidance instead.

  • MiniMaxAI/MiniMax-H3 weights
  • Diffusers pipeline
  • Community license terms

What Makes MiniMax H3 Different

H3 unifies jobs that used to need a relay of separate models, and it ships weights you can download - a combination few video models offer.

One model for generation, referencing, and editing

Text, images, video, and audio enter one shared context. Character identity from a photo, camera language from a clip, a voice from a recording - H3 reads them together instead of handing off between specialists.

Native synchronized stereo audio

Every generation carries its own soundtrack: ambience, foley, and speech generated alongside the picture, with stable dialogue support for 11 languages including English and Chinese.

Readable on-screen text

H3 renders legible type, which puts title cards, signage, and brand marks inside a prompt's reach. MiniMax's launch write-up calls out accurate text and brand rendering as an early-testing strength.

Open weights, competitive pricing

The complete weights are published on Hugging Face to support further development, including fine-tuning. MiniMax also positions 2K output at less than a third of mainstream per-second pricing; on fal the 2K rate is $0.26 per second.

How to Generate with MiniMax H3

The workflow below is documentation-based, following MiniMax's model card and published prompting guidance - pick the input mode, assemble labeled references, write the brief, then scale resolution.

1

Pick the input mode that matches your materials

No media is text-to-video. One or two frame images (opening or closing frame) is the FL2VA mode. Identity, style, motion, voice, or editing references use the Ref2VA mode.

2

Assemble references and label their jobs

Ref2VA accepts up to 9 images, 3 video clips of 2-15 seconds each, and 3 audio clips, capped at 12 files. Cite each one as Image 1, Video 1, Audio 1 - input order is semantic, and audio must travel with at least one image or video.

3

Write the shot as a timed brief

Structure the prompt in timed blocks with one primary change per beat, an explicit camera move, a directed audio track, and stated negatives. Then render at 768p and run the 2K regeneration pass for the final master.

MiniMax H3 Prompt Tips

H3 follows direction better than description. These four habits come straight from the published prompting guidance and separate usable clips from mush.

Assign a job to every reference

Never send unlabeled media. Write "Use Image 1 as the locked identity, Video 1 for the camera move, Audio 1 as the voice" - named jobs bind references to roles; unlabeled inputs leave the model guessing.

Time the shot in blocks

Mark beats as [0 to 4 seconds], [4 to 8 seconds], and so on. Give each beat one primary change and an observable end state, and put the most important beat in the middle of the timeline - the final beat is the most likely to get squeezed.

Direct the audio as its own track

Name room tone, foley, and music placement explicitly. Tag speakers (S1, S2), put verbatim dialogue in a [Language] tag from the 11 supported languages, and describe when lips close.

State negatives and what must not change

"No soft dissolves, no camera shake, wardrobe identical in every shot" - negative direction lands unusually well. Name the invariant details (hair, closures, footwear) and describe transitions as physical events.

MiniMax H3 FAQ

Straight answers about H3's weights, outputs, inputs, pricing, and how it relates to AI Fruit.

The weights are published on Hugging Face under the MiniMax H3 Community License, which is a vendor community license rather than an OSI open-source license. The open release includes the base generator and the 2K regeneration pass with a Diffusers pipeline; the hosted context-processing system that precedes generation is not part of the open release, so local runs rely on the published prompting guidance.

Per the model card: clips of 4 to 15 seconds at 24 FPS, a default shorter side of 768 pixels with 2K available through the H3-Regenerate-2K pass, 32 kHz stereo audio generated natively with the picture, and a wide set of aspect ratios including 21:9, 16:9, 4:3, 1:1, 3:4, and 9:16.

Text always works alone. The first-and-last-frame mode takes zero, one, or two images. The omni-reference mode takes up to 9 images, 3 video clips of 2-15 seconds each (15 seconds total), and 3 audio clips, capped at 12 files; audio must be accompanied by at least one image or video.

Try the Effect on AI Fruit's Engine

H3 itself runs on MiniMax's app, fal, or its open weights - but the prompt style works anywhere. Open the generator below, paste an H3-style brief, and generate with an AI Fruit model you can pick and price before spending credits.

Open the Generator