MiniMax H3 Video Generation Cut To 31.6 Seconds On A Single RTX 5090
If a single user can reach 6.12x on a 19.5GB open model at home, the argument for waiting on the unreleased H3 Max weights tilts toward prompt adherence rather than speed.
Reporting from 1 source: GameBusiness.jp.
A first-person report on GameBusiness.jp timed an RTX 5090 running MiniMax H3 locally. One 5-second i2v clip went from 193.5 seconds to 31.6 seconds in 41 days, a 6.12x gain, using a 20-step baseline, the official Turbo 8-step LoRA, Hao AI Lab's FastH3 4-step, a community-converted TaoMate 3-step LoRA, INT8 attention, and sparse attention. H3 Max weights have not been released 28 days after fal said on X that it would publish them. Cloud H3 Max Turbo ran at 1.49 seconds on different hardware.
The baseline is 1280x704, 5.2 seconds, 124 frames, same first image, same prompt, same three seeds, i2v, four runs per condition with the first discarded and the median of three kept. Each LoRA arrived on a different date, which is why the gains landed in bursts rather than a curve.
Three step-plus-attention combinations carried most of it: the official Turbo 8-step LoRA reached 90.7 seconds, FastH3 4-step reached 55.3 seconds, and a community ComfyUI conversion of a TaoMate 3-step LoRA reached 44.1 seconds. Dropping INT8 attention from the last of those fell back to 59.0 seconds, so the quantization alone was worth 1.34x. Sparse attention at 0.3 then 0.1 brought the clip to 35.0 and 31.6 seconds.
The next bottleneck is VAE decoding, at 39 percent of total time. ComfyUI 0.36 shipped a new VAE; tested in the production setup, it came out even. A 30-clip music video moved from 1 hour 37 minutes to 16 minutes.
Synthesized by Yomimono from the 1 cited source below, including Japanese-language reporting where cited, then editorially reviewed before publishing.