comfyui minimax h3
Craft clips with matched stereo audio through the comfyui minimax h3 workflow
AI Video Prompt Generator
10s

Feedback

AI Ad Video Example

Loading...

comfyui minimax h3

Create open-weight video inside ComfyUI with MiniMax H3 — text-to-video, image-to-video, and reference-to-video modes with native stereo sound at up to 2K 24fps.

All Tools

Discover our comprehensive AI-powered animation toolkit

What Makes the comfyui minimax h3 Pipeline Stand Out

This comfyui minimax h3 pipeline unlocks MiniMax's general-purpose omni-modal generation model as open weights inside ComfyUI. It processes text, images, video, and audio within one shared context, delivering clips with native stereo audio — dialogue, effects, and music generated collectively in a single pass. You get up to 2K at 24fps for roughly 15 seconds, with node-level control over every setting.

  • Synced Stereo Sound
    Speech, sound effects, and music render together with the footage in one MP4, perfectly synchronized by the comfyui minimax h3 pipeline in a single forward pass.
  • Full Local Control
    Keep the comfyui minimax h3 model on your own machine and fine-tune resolution, duration, and every diffusion setting — no API restrictions or hidden limits.
  • Multi-Source Reference Input
    Blend text, images, video, and audio references in a single generation to preserve a character, style, motion, camera movement, or voice through the comfyui minimax h3 nodes.

Running the comfyui minimax h3 Workflow in 3 Steps

Produce open-weight videos with built-in audio using the comfyui minimax h3 workflow — just three steps from setup to export.

Core Capabilities of the comfyui minimax h3 Workflow

From three native ComfyUI templates and open-weight multimodal generation to synced stereo audio, reference-driven control, and optional Sage Attention acceleration — the comfyui minimax h3 workflow forms a complete local video production toolkit.

Three Ready-Made Templates

The comfyui minimax h3 template pack includes text-to-video, image-to-video, and reference-to-video examples, each designed for one generation mode out of the box.

Unified Multimodal Context

The comfyui minimax h3 model interprets text, images, video, and audio together in a single context, letting you combine all reference types in one generation.

Reference-First Creation

Preserve a character's identity, an artistic style, a motion, a camera move, or a voice from your references — supporting up to 9 images, 3 videos, and 3 audio clips through the comfyui minimax h3 R2V node.

Clear Text & Brand Display

The comfyui minimax h3 model renders spelled-out words and brand assets accurately, following instructions that describe reference relationships in natural language.

Sage Attention Acceleration

Nearly double your generation speed with little quality loss by inserting the Patch Sage Attention KJ node into the comfyui minimax h3 flow.

Flexible Resolution & Duration Matrix

The comfyui minimax h3 Resolution Selector derives width and height from aspect ratio plus megapixels, snapped to the model's 32-multiple canvas and 17-frame-per-block duration at 24fps.

FAQ

comfyui minimax h3 — Common Questions

Straight answers about using the MiniMax H3 model inside ComfyUI and getting the most from your setup.

1

What does the comfyui minimax h3 workflow do?

It is ComfyUI's native integration of MiniMax H3, MiniMax's general-purpose omni-modal generation model released as open weights. The flow creates video with native stereo audio from text, images, video, and audio references in a single forward pass.

2

What quality and length can I expect?

The comfyui minimax h3 workflow produces up to 2K resolution at 24fps for roughly 15 seconds. Its native canvas starts from a 768px short edge, capped at 768x1344 pixels and rounded to multiples of 32.

3

Which generation modes are available?

The comfyui minimax h3 template library includes three examples: text-to-video (T2V), image-to-video (I2V) with optional first/last-frame control, and reference-to-video (R2V) that locks in character, style, motion, camera, or voice.

4

Does the output include audio?

Absolutely — the comfyui minimax h3 model generates native stereo audio covering voice, sound effects, and music, all modeled with the video in one pass and delivered in a single MP4 file.

5

How do I begin with this workflow?

Upgrade ComfyUI to version 0.30.0 or later, head to Template Library > Video, select a comfyui minimax h3 flow, and follow the pop-up to download models from the Hugging Face Comfy-Org/MiniMax-H3 repository.

6

Is there a way to speed it up?

Yes — install SageAttention and the KJNodes custom nodes, then add a Patch Sage Attention KJ node between the UNETLoader and BasicGuider in the comfyui minimax h3 workflow to roughly double its generation speed.

Jump Into the comfyui minimax h3 Workflow Today

Run MiniMax H3 in ComfyUI with open weights, native stereo audio, and complete parameter control — text-to-video, image-to-video, and reference-to-video flows are ready when you are.