ComfyUI · NSFW · 18+
MiniMax-H3 NSFW video workflow
The Video Pack runs on MiniMax-H3, the open image-to-video model by MiniMax, packaged for ComfyUI through Comfy-Org on Hugging Face. The pack ships the pruned int8 reference-to-video build, the Qwen3-VL 32B text encoder, separate video and audio VAEs, and the ref2v turbo LoRA. A source image and a prompt produce a clip of 5 to 15 seconds whose soundtrack matches the motion.
What the pack ships
| Component | File | Notes |
|---|---|---|
| Diffusion model | Minimax-h3 ref2va, pruned, int8 | 21 GB, needs a 32 GB card |
| Text encoder | Qwen3-VL 32B (nvfp4 awq) | Runs on Ampere up, no Blackwell required |
| Video VAE | minimax_h3_video_vae, int8 | Encodes the frames |
| Audio VAE | minimax_h3_audio_vae, fp32 | Generates the soundtrack |
| Turbo LoRA | minimax_h3_ref2v_turbo_4step | Ships applied, speeds sampling |
| Tuned LoRA | 1 quality LoRA from our Civitai | Locks in the final look |
Everything installs with one command. Sampler, CFG, and LoRA settings arrive pre-calibrated. Leave them as they ship.
How a clip gets made
- The source image becomes the first frame of the clip.
- The diffusion model animates it, 24 fps, 5 to 15 seconds.
- The audio VAE generates the soundtrack in the same pass, so sound matches motion.
- A two-stage quality pipeline renders, then refines the result.
The prompt carries two blocks: one for the picture, one for the sound. Describe one continuous shot in temporal order: what moves, then the lighting and camera, then the matching sounds.
Writing prompts that hold up
- Describe motion, not appearance. The source image already sets the look.
- One primary action per clip. Layered actions lower coherence.
- Do not contradict the source image. The clip starts from it.
- One continuous action scales better than a sequence of position changes.
Hardware
NVIDIA GPU with 32 GB+ VRAM (RTX 5090). On RunPod, RunPod RTX 5090 / RTX 5000 Ada / A6000+. Disk needs 60 GB for models. Setup guides are linked below.
Adults only (18+). All sample output is synthetic, computer-generated fiction. Last verified October 2026.