Feedback
AI Ad Video Example
Loading...
comfyui minimax h3
Create audio-rich clips using the comfyui minimax h3 workflow—MiniMax H3 open weights turn text, images, or references into stereo-sound video at 2K/24fps.
All Tools
Discover our comprehensive AI-powered animation toolkit
MiniMax H3
MiniMax H3 AI Video Generator
Seedance 2.5
The Future of AI Video Is Here.

Seedance 2.0
The Future of AI Video Is Here.

Veo3.1
Create Stunning Videos with Veo3.1

Kling 3.0
Next-Gen AI Video Generator
Grok Video Generator
Create Videos from Text or Images with AI
MiniMax H3 video generator
MiniMax H3 AI Video Generator

Seedance 2.0
The Future of AI Video Is Here.
Why Creators Choose the comfyui minimax h3 Workflow
Running MiniMax H3's open-weight, omni-modal model inside ComfyUI, the comfyui minimax h3 workflow lets you feed text, visuals, motion clips, and sound into one context. In a single forward pass, it produces an MP4 with dialogue, effects, and music already embedded—up to 2K/24fps and around 15 seconds—while leaving every node adjustable.
- Stereo Sound, Synced by DefaultVoice, sound effects, and music are generated alongside the visuals in one MP4—so the comfyui minimax h3 output is always audio-aligned and ready to play.
- Open-Weight FlexibilityBecause MiniMax H3 ships as open weights, you operate the comfyui minimax h3 nodes on your own hardware and fine-tune resolution, length, and diffusion settings without API quotas.
- Reference Anything, Combine EverythingUse any mix of text prompts, stills, footage, or sound clips to steer identity, look, movement, framing, or voice—each reference is routed through the comfyui minimax h3 node graph.
Getting Started with the comfyui minimax h3 Workflow
Follow this quick guide to create audio-synced, open-weight videos using the comfyui minimax h3 workflow.
Built-In Capabilities of the comfyui minimax h3 Workflow
From three ready-made ComfyUI templates and open-weight omni-modal generation to stereo audio, reference-guided creation, and Sage Attention boosts—the comfyui minimax h3 pipeline covers your entire local video workflow.
Three Ready-to-Run Template Types
The comfyui minimax h3 template pack includes T2V, I2V, and R2V examples, so each prompting style has its own ready-to-run ComfyUI graph.
Unified Multimodal Understanding
MiniMax H3 processes text, stills, motion, and sound in one unified context, letting the comfyui minimax h3 node set fuse every input type into a single generation.
Reference-Driven Result Lock-in
Anchor a character, visual style, action, camera angle, or vocal timbre from provided assets—up to nine images, three video clips, and three audio files through the comfyui minimax h3 R2V node.
Text and Logo Friendly Rendering
On-screen words and logos come out crisp and correct with the comfyui minimax h3 model, and natural-language instructions accurately describe how references relate to each other.
Sage Attention for 2x Speed
Add the Patch Sage Attention KJ node to the comfyui minimax h3 graph to cut render time by about half while doing little damage to output quality.
Precision Resolution & Duration Sizing
With the comfyui minimax h3 resolution selector, width and height are derived from your aspect ratio and megapixel target, then snapped to the 32-step grid and 17-frame-per-block rule at 24fps.
Frequently Asked Questions About the comfyui minimax h3 Workflow
Quick answers on installing, generating, and tuning the MiniMax H3 model through the comfyui minimax h3 workflow.
What does the comfyui minimax h3 workflow actually do?
It's ComfyUI's built-in integration for MiniMax H3, an open-weight, omni-modal model from MiniMax. In one pass, it turns text, images, clips, and sound references into video that includes stereo audio.
What resolutions and frame rates can I expect?
The comfyui minimax h3 workflow supports clips up to 15 seconds long at 2K/24fps. Its internal canvas starts with a 768px short edge, never exceeds 768x1344 pixels, and snaps to multiples of 32.
Are different prompt modes available?
Yes. The comfyui minimax h3 library comes with three ready examples: T2V, I2V with optional start/end frame control, and R2V that locks a character, style, movement, shot, or voice.
Can it really create sound along with video?
Absolutely—the comfyui minimax h3 model outputs stereo soundtracks with speech, effects, and score, rendered together with the visuals and embedded in one MP4.
How do I set up the workflow for the first time?
Upgrade ComfyUI to 0.30.0 or newer, go to Template Library > Video, select a comfyui minimax h3 template, and follow the prompt to fetch models from the Hugging Face Comfy-Org/MiniMax-H3 repo.
Is there a way to make the workflow run faster?
Yes. Install SageAttention and KJNodes, then place a Patch Sage Attention KJ node between the UNETLoader and BasicGuider in the comfyui minimax h3 graph to nearly double render speed.
Jump Into the comfyui minimax h3 Workflow Today
Launch the comfyui minimax h3 workflow on your own ComfyUI setup and create open-weight, stereo-sound clips from text, images, or references—with every parameter under your control.
