MiniMax H3 Video AI Ranks Second in Blind Test, Open Weights Coming
MiniMax H3's second-place blind ranking and upcoming open-weight release position it as a strong open alternative in a field dominated by closed models.
Reporting from 1 source: GIGAZINE.
Chinese AI firm MiniMax announced MiniMax H3, a video generation model that creates up to 15 seconds of audio-synced video from text, images, or voice. It supports 768p or 2K resolution and accepts multiple media inputs at once. In a blind test by Artificial Analysis, H3 ranked second for text-to-video, behind Gemini Omni Flash and ahead of Seedance 2.0. The model is available via API now, with open weights planned soon.
MiniMax, the Chinese AI company behind the open-weight M3 language model, has unveiled its latest video generator. MiniMax H3 accepts text, video, image, and audio inputs, and can combine several media types in a single prompt. It generates clips up to 15 seconds long at 768p or 2K resolution.
Independent testing firm Artificial Analysis ran a blind comparison where models generated videos from identical prompts without names attached. For text-to-video with audio, H3 placed second, trailing Gemini Omni Flash but ahead of Seedance 2.0. For image-to-video with audio, H3 took third behind Seedance 2.0 and Gemini Omni Flash.
The model is already usable through MiniMax's API, and third-party tools like ComfyUI can call it. MiniMax says the model itself will be released as open weights in the near future.
Synthesized by Yomimono from the 1 cited source below, including Japanese-language reporting where cited, then editorially reviewed before publishing.