ComfyUI Node

MiniMaxH3Director

A real timeline for MiniMax H3, without leaving ComfyUI

By AIMixer·Created 2 months ago·Updated about 14 hours ago· 2,066
MiniMaxH3Director
  • model
  • video_vae
  • audio_vae
  • clip
  • bd_grp_sample
  • i2v_groups
  • r2v_groups
  • semantic_bridge
  • selflift
  • refine
  • face_refine
  • bd_grp_advanced
  • bd_grp_perf
  • sigmas
  • images
  • audio
  • fps
  • frame_count
  • source_images
  • report
  • images_pre_refine
  • images_pre_face_refine
◄task_typet2v — 文生视频(Text to Video)►
◄global_promptA cinematic scene with natural motion and synchronized ambience►
◄cfg1.00►
◄seed0►
◄frame_rate24.00►
◄width864►
◄height480►
◄ref_max_size864►
◄total_frames124►
◄timeline_data►
◄steps25►
◄samplerres_multistep►
◄schedulersimple►
◄shift_video12.00►
◄shift_audio3.00►
◄clear_vram_between_segmentstrue►
◄clear_vram_before_refinefalse►
◄clear_vram_before_face_refinefalse►
◄cache_frames_codecraw►
◄export_source_imagesfalse►
◄export_pre_face_refinefalse►

MiniMax H3 is the video model that dropped in late July 2026 and immediately had the local community clearing hard-drive space: open weights, up to 15 seconds at 2K, native stereo audio, and reference-to-video input that ComfyUI supports natively. Native support, though, means native nodes - and the native MiniMaxH3ImageToVideo / MiniMaxH3ReferenceToVideo nodes are bare function calls. If you want to compose a scene from multiple references, trim clips, order shots, and manage prompts without a spreadsheet, you need something on top.

MiniMax H3 Director is that something. It's a timeline-based authoring node - a direct descendant of the LTX Director idea that WhatDreamsCost made famous - rebuilt for H3. One node holds your images, videos, audio, trims, ordering, and prompts, and hands the whole thing to ComfyUI's native H3 implementation.

The mechanism

The Director doesn't reimplement MiniMax H3 - the README is explicit that it routes to ComfyUI's built-in nodes. It does the authoring: you drop media into two timeline lanes (Image/Video and Audio) via drag-and-drop, paste, or file picker, set per-clip trim ranges, reorder tiles by dragging, and write your prompt in a mode-specific builder. On queue it decodes everything (video cropped to your trim, audio extracted and cropped the same way), validates H3's hard limits, and emits a structured guide dictionary plus a resolved prompt to the sibling MiniMax H3 Director Guide node, which makes the actual native call.

Two modes:

  • FL2VA - text-to-video, first-frame, or first+last-frame interpolation. Up to 2 image slots; audio/video blocked. The builder has guided fields (integrated_multimodal_description, overall_soundscape, non_diegetic_music) and auto-inserts the alignment lines that H3 wants.
  • REF2VA - the reference mode that makes H3 interesting. Up to 9 images, 3 videos, 3 audio clips (12 files total), each video switchable V / A / V+A. The builder here is six free-text sections following H3's official full-reference format, with helper buttons: Insert [Shot N], Prefill Labels & Summary, and Preview Prompt so you see the exact assembled prompt before you commit.

Inputs that matter

  • mode - T2VA / I2VA / FL2VA / L2VA / REF2VA. Everything about the node's behavior follows this switch.
  • duration (default 5s) and frame_rate (default 24) - set your clip length and FPS. frame_rate is also emitted as an output so downstream nodes read the effective value.
  • The Resolution panel (Aspect / Resolution / Input scaling, all default Auto) drives the output canvas on H3's 16px grid. Auto resolution sets a 768px short side; Auto aspect follows your first visual reference.
  • The optional fl2va_model / ref2va_model sockets enable lazy loading - only the model for your active mode gets requested.
  • external_prompt_overwrite (STRING) and the paired external_width_overwrite / external_height_overwrite (INT) let you bypass the builder and canvas entirely when you already know exactly what you want.

The outputs

The one that matters is guide - it feeds the Director Guide node. Also useful: duration, frame_rate, width, height, positive_prompt (the assembled text, handy for debugging), and model / fl2va_requested / ref2va_requested for routing.

Installing it

ComfyUI Manager (search DaSiWa-Nodes), or:

cd ComfyUI/custom_nodes
git clone https://github.com/darksidewalker/ComfyUI-DaSiWa-Nodes
pip install -r requirements.txt

then restart. You also need a ComfyUI build with native MiniMax H3 support and the H3 model files (checkpoint, Qwen3-VL text encoder, visual VAE, and the audio VAE for REF2VA) via ComfyUI's model manager. This is where the pack's RTX requirement gets fuzzy - H3 itself runs on consumer cards (it reportedly runs on a 3060, slowly), so this node needs no NVIDIA SDK, just the model and patience.

Where people get burned

  • Forgetting the audio VAE in REF2VA. The Guide refuses to run without audio_vae connected. Wire vae/minimax_h3_audio_vae_fp32.safetensors in.
  • Hitting the limit walls. References must be 2–15 seconds, combined visual ≤15s, combined audio ≤15s. The node validates this and shows red status messages - believe them.
  • 16px vs 32px grid trap. H3's VAE works on 16px latent cells, but its transformer groups them in 2×2 patches - so a width/height that's a clean multiple of 16 but not 32 can fail at sampling. Use the resolution panel's presets rather than raw CUSTOM values unless you know what you're doing.
  • Media that stays on your disk. Files you upload get stored under ComfyUI's input/ directory, referenced by relative path. Delete them there and old timelines break.

This is a big, opinionated node with a real learning curve, and it's the reason to install the pack if you do H3 video. The one thing it can't do is make a 15-second 2K generation fast - that's H3's problem, not the Director's.

CategoryMiniMaxH3

Inputs (35)

NameTypeDefaultDescription
modelMODELMiniMax H3 UNET (UNETLoader).
video_vaeVAEMiniMax H3 video VAE (minimax_h3_video_vae).
audio_vaeVAEMiniMax H3 audio VAE (minimax_h3_audio_vae). Required for r2v / v2v / rv2v.
clipCLIPCLIPLoader type=minimax (qwen3vl).
task_typeCOMBOt2v — 文生视频(Text to Video)MiniMax H3 支持 t2v / i2v / fl2v / r2v / v2v / rv2v / mixed。提示词直接送入 MiniMaxH3ImageToVideo 或 MiniMaxH3ReferenceToVideo(内部 tokenize)。mixed 为每段自选 t2v/i2v/fl2v/r2v;r2v 用 <Picture 1>;v2v/rv2v 为源视频时间轴编辑(自动绑定 <Video 1>);rv2v 另可挂参考图。
global_promptSTRINGA cinematic scene with natural motion and synchronized ambienceUser prompt — sent directly to MiniMaxH3ImageToVideo / ReferenceToVideo. r2v: <Picture 1>. v2v: source-timeline edit (<Video 1>). rv2v: source timeline + reference images (<Video 1> + <Picture N>).
bd_grp_sampleBDGROUP采样设置—
cfgFLOAT1.000–30CFG for KSampler.
seedINT00–18446744073709550000Random seed for sampling.
frame_rateFLOAT24.001–240Timeline / output FPS (H3 trained at 24).
widthINT86432–8192—
heightINT48032–8192—
ref_max_sizeINT86432–8192—
total_framesINT1245–100000Frame count at 24 fps; snapped to MiniMax 17k+5 grid (124 ≈ 5s).
timeline_dataSTRINGInternal — video, segments, refs (populated by UI).
i2v_groupsoptMMX_DIR_GROUPExternal Image to Video group(s) (t2v / i2v / fl2v). When connected, overrides UI cards for execution (external priority). Connect Group (Image to Video).group, or Groups Combine.
r2v_groupsoptMMX_DIR_GROUPExternal Reference to Video group(s). When connected, overrides UI cards for execution (external priority). Connect Group (Reference to Video).group, or Groups Combine.
semantic_bridgeoptMMX_DIR_SEMANTIC_BRIDGEOptional Semantic Bridge node (above SelfLift). When connected, official cond tokens are rewritten with the student MLP (RMS-norm → residual mix). Unconnected = identical. Distilled on FL2VA; r2v / v2v / rv2v is forced-compat — use with care.
selfliftoptMMX_DIR_SELFLIFTOptional SelfLift node (above Refine). When connected, first-pass is low-res prefix + 3D lift + high-res tail on this Director canvas. Unconnected = current single-stage sample. Refine may still upscale afterward (e.g. 1.0MP first pass → 2.0MP). Timeline continuity keeps native low-res carry + high-res pin. Euler only.
refineoptMMX_DIR_REFINEOptional Refine node. When connected, each segment runs a second sample pass (same-size refine, or upscale then sample). Wire a MODEL into Refine.refine_model to use a different UNET for that pass; unwired uses this Director model. images is the refined result; images_pre_refine is the first pass. Unconnected = single-pass (current behavior).
face_refineoptMMX_DIR_FACE_REFINEOptional FaceRefine node. When connected, Director tracks the face on the final decoded frames (after Refine if that is also wired), re-samples the crop, and pastes the face back. images is after stitch; images_pre_face_refine is before stitch when「输出修脸前」is on (otherwise that output is blocked). Unconnected = no face pass (that output stays blocked).
bd_grp_advancedoptBDGROUP高级采样—
stepsoptINT251–200一采步数(官方模板 25)。接了 sigmas 口后忽略此项,改用外接噪声表。
sampleroptCOMBOres_multistepOfficial template: KSamplerSelect res_multistep.
scheduleroptCOMBOsimple一采调度器(官方模板 simple)。接了 sigmas 口后忽略此项,改用外接噪声表。
shift_videooptFLOAT12.000.01–100MiniMaxH3SigmaShift shift_video.
shift_audiooptFLOAT3.000.01–100MiniMaxH3SigmaShift shift_audio.
bd_grp_perfoptBDGROUP性能—
clear_vram_between_segmentsoptBOOLEANtrue段间清理显存:每段结束后卸载模型并清空 CUDA 缓存。
clear_vram_before_refineoptBOOLEANfalse二采前清理显存:一采结束后、放大或二采开始前卸载模型并清空 CUDA 缓存。默认关。24GB 或一采/二采不同 UNET 时勾上,可降低二采峰值,但每段会多一次加载。
clear_vram_before_face_refineoptBOOLEANfalse脸修前清理显存:成片解码后、FaceRefine 开始前卸载模型并清空 CUDA 缓存。默认关。未接 FaceRefine 时无效。24GB 或解码后立刻 OOM 时勾上,但每段会多一次加载。
cache_frames_codecoptCOMBOraw分段缓存的像素帧怎么存。raw:uint8 .pt,读写快,磁盘大。ffv1:无损压缩,体积大约三分之一,命中和写入都会多一次编解码。两种都能读;某一段被重新写入时才换成当前选项。latent / 音频缓存不受影响。
export_source_imagesoptBOOLEANfalse将时间轴原片解码到独立的 source_images 输出口;需将 source_images 另接预览/合成节点才能查看,不会改变主 images。默认关以节省内存。
export_pre_face_refineoptBOOLEANfalse将修脸前的视频输出到 images_pre_face_refine,方便和 images 对比。分段导出时同时写入 seg_XXXX_facepre.mp4。默认关:该口阻断、下游不执行,也不占成片内存。未接 FaceRefine 时无效。
sigmasoptSIGMAS可选。一采噪声表,接 BasicScheduler 或 ManualSigmas。接线后覆盖导演台「步数」和「调度器」(采样器下拉仍有效)。BasicScheduler 请接 SigmaShift 之后的同一套 H3 MODEL。不接则仍用步数 + 调度器、denoise=1 自动算表。

Outputs (8)

NameTypeDescription
imagesIMAGE—
audioAUDIO—
fpsFLOAT—
frame_countINT—
source_imagesIMAGE—
reportSTRING—
images_pre_refineIMAGE—
images_pre_face_refineIMAGE—