🎬 MiniMax-H3 · Text-to-Video

Text-to-video with synchronized stereo audio from MiniMaxAI/MiniMax-H3 — a 33B omni-modal generator run here as a pruned NVFP4 transformer with a local layer-50 Qwen3-VL conditioner.

Engine status: Ready · NVFP4 · Cache-DiT F1B0+TaylorSeer 0.08 / forecast 0.65 · pruned AdaLN curve · fused QKV/QK-norm/RoPE · lilcheaty/MiniMax-H3-NVFP4 · VAEs full precision · placement lazy · attention _native_cudnn · loaded in 136s · conditioner local layer-50 Qwen3-VL NVFP4-AWQ weights / BF16 GEMM · Comfy-Org/MiniMax-H3 · TAE live previews

Speed & quality preset
Balanced — best overall (recommended): Full quality schedule with conservative Cache-DiT acceleration. | Turbo 8-step — faster, cleaner: Distilled eight-step path with better Turbo consistency. | Turbo 4-step — fastest, more artifacts: LightX2V v0.1 four-step preview; maximum speed with some detail loss. | Exact 28-step — maximum fidelity: Dense reference path with approximate caching disabled. | Ultra cache — experimental speed: Aggressive forecasting on the full schedule; inspect results carefully. | Custom — manual controls: Expose schedule, cache engine, and Turbo LoRA controls.
Canvas
2 14
Example prompts