Nodes/ComfyUI-MinimaxH3DYTsc/minimaxH3DYTsc6.0 Director
ComfyUI Node

minimaxH3DYTsc6.0 Director

An entire MiniMax pipeline hiding behind one node

By 792877530-star·Created 6 days ago·Updated 6 days ago· 2
minimaxH3DYTsc6.0 Director
  • bd_grp_sample
  • bd_grp_advanced
  • bd_grp_perf
  • model
  • video_vae
  • audio_vae
  • clip
  • model_ref
  • ltx_model
  • ltx_clip
  • ltx_vae
  • ltx_upscaler
  • ltx_iclora_model
  • images
  • audio
  • fps
  • compare_p1
  • compare_p2
◄task_typer2v — 参考素材生视频(Reference to Video)►
◄global_promptA cinematic scene with natural motion and synchronized ambience►
◄cfg1.00►
◄seed0►
◄frame_rate24.00►
◄width864►
◄height480►
◄ref_max_size864►
◄total_frames124►
◄timeline_data►
◄steps25►
◄samplerres_multistep►
◄schedulersimple►
◄shift_video12.00►
◄shift_audio3.00►
◄clear_vram_between_segmentstrue►
◄export_source_imagesfalse►
◄preview_taetrue►
◄preview_overhead_percent10►
◄sag_enabledtrue►
◄ltx_iclora_downscale1►

One node, a whole studio

minimaxH3DYTsc6.0 Director (class MinimaxH3DYTScDirector) is the reason this pack exists. It is not a loader or a sampler - it's conditioning, sampling, decode, optional refinement and saving, all inside a single node with its own timeline UI drawn on the node body. You give it a prompt and a frame count and press run.

That's a fork in the road. If you like wiring graphs, the pack's MinimaxH3DYTScConditioning node plus ComfyUI's official H3 nodes does everything this does, visibly. The Director is for shots in a list - several prompts, several durations, reference material per shot, continuation from the previous segment's last frame - without 200 nodes of canvas.

The task_type enum has two entries, and they lead to genuinely different machinery. t2i is a text-to-image workspace that does not touch H3 at all: it runs a Krea2 Turbo pipeline with its own model card, and H3 checkpoints are ignored there. r2v is the video side, and despite the name it's a unified workspace - no reference material means plain text-to-video, keyframes mean i2v/fl2v, and reference images/videos/audio route through H3's reference conditioning with continuation between segments.

How it works

Per segment it builds the same joint audio+video latent the official nodes do, samples it in a single stage (KSampler plus MiniMaxH3SigmaShift), and decodes through LTXVSeparateAVLatent, which splits the joint latent into the video and audio halves. Then comes an optional second pass - 二采 - either a low-strength H3 resample or the LTX 2.5 two-stage refine (euler_ancestral, CFG 1, a fixed low-noise sigma tail, first-frame guidance, audio latent untouched). It has a quality gate: if the second pass doesn't measurably sharpen, you get the first-pass result instead of mush.

That's what the two comparison outputs are: compare_p1 is the first pass, compare_p2 is the refined/final one. With refinement off, they're the same frames.

Model loading is internal too: a card inside the node points at your UNET, reference UNET, CLIP, video VAE, audio VAE, Turbo LoRAs and the optional LTX 2.5 refine stack. Prefer canvas loaders? The optional model, clip, video_vae, audio_vae, model_ref and ltx_* inputs override the card when connected.

The inputs that matter

Beyond task_type and global_prompt: total_frames defaults to 124 - about five seconds at 24 fps - and snaps to H3's 17k+5 grid. width/height default to 864×480, frame_rate to 24 because that's what H3 was trained at, steps to 25, cfg to 1 (the official template's value; don't raise it), sampler to res_multistep and scheduler to simple. shift_video (12) and shift_audio (3) are the SigmaShift values - leave them unless you're chasing a specific look.

Then the performance toggles, which are where the real decisions live. clear_vram_between_segments unloads between segments - keep it on unless you have VRAM to spare. sag_enabled switches attention to Comfy Kitchen's INT8 path, which the author notes is close to SageAttention in speed and memory and actually works on H3; SageAttention's int8 route is disabled on purpose because it collapses on H3's long sequences. preview_tae gives you real-colour frames mid-sampling, but only if taeh3.safetensors is in models/vae_approx - otherwise it silently degrades to the cheap latent2rgb preview, which is why people think the preview looks bad.

timeline_data is the internal string the UI writes. Don't hand-edit it.

Outputs: images (the finished frames, as a list), audio, fps, plus compare_p1 and compare_p2. The node is also an output node - it saves the muxed video itself through the same path as ComfyUI's CreateVideo + SaveVideo, so you can delete those nodes from your graph. Wire compare_p2 (and compare_p1), images, audio and fps into the pack's MinimaxH3DYTScVideoCompare if you want the before/after view.

Install

Same pack as the rest of these nodes:

cd ComfyUI/custom_nodes
git clone https://github.com/792877530-star/ComfyUI-MinimaxH3DYTsc

Restart ComfyUI. Manager users: search minimaxH3DYTsc6.0. You need a recent ComfyUI (the pack calls comfy_extras.nodes_minimax_h3 and checks its signatures), the H3 UNETs, the Qwen3-VL H3 text encoder loaded as type minimax, the video VAE and minimax_h3_audio_vae_fp32.safetensors. The requirements are opencv-python-headless, imageio-ffmpeg, scenedetect and comfy-kitchen>=0.2.34. The LTX 2.5 refine stack is optional - skip the models and the second pass just doesn't happen.

Where people get burned

  • The model card looks filled in and nothing loads. Every default slot ships with a filename but as an open circuit - you have to tick it to actually load. The pack even detects the "everything disconnected" state, restores the slots by filename and warns you, because that state can't generate anything.
  • /minimax/director/* returns 404. The node's HTTP routes register lazily; if PromptServer wasn't up yet at plugin load you get a warning in the console and a broken panel. Restart ComfyUI.
  • A reference segment dies on the Audio VAE. Same rule as everywhere in this pack: reference conditioning needs the audio VAE loaded in the model card.
  • This is a UI-heavy node, so the ComfyUI frontend can break it. Nodes 2.0 - the Vue rewrite of the canvas - has been rough on packs that draw their own panels, and the lesson from rgthree is that a single maintainer may not port. If the timeline panel misbehaves, the legacy canvas is still the safe fallback.
  • GGUF checkpoints need ComfyUI-GGUF installed and a restart. And many of this pack's error strings are in Chinese: accurate, but not searchable.
CategoryMiniMaxH3

Inputs (34)

NameTypeDefaultDescription
task_typeCOMBOr2v — 参考素材生视频(Reference to Video)MiniMax H3 导演台保留独立文生图与参考素材生视频。参考素材生视频统一承载无参考生成、首帧、尾帧、中间帧、图片/视频/音频参考,并支持分段续接上一段生成结果。
global_promptSTRINGA cinematic scene with natural motion and synchronized ambienceUser prompt — sent directly to MiniMaxH3ImageToVideo / ReferenceToVideo. r2v: <Picture 1>. v2v: source-timeline edit (<Video 1>). rv2v: source timeline + reference images (<Video 1> + <Picture N>).
bd_grp_sampleBDGROUP采样设置—
cfgFLOAT1.000–30CFG for KSampler.
seedINT00–18446744073709550000Random seed for sampling.
frame_rateFLOAT24.001–240Timeline / output FPS (H3 trained at 24).
widthINT86432–8192—
heightINT48032–8192—
ref_max_sizeINT86432–8192—
total_framesINT1245–100000Frame count at 24 fps; snapped to MiniMax 17k+5 grid (124 ≈ 5s).
timeline_dataSTRINGInternal — video, segments, refs (populated by UI).
bd_grp_advancedoptBDGROUP高级采样—
stepsoptINT251–200Sampling steps — official template: 25.
sampleroptCOMBOres_multistepOfficial template: KSamplerSelect res_multistep.
scheduleroptCOMBOsimpleOfficial template: BasicScheduler simple.
shift_videooptFLOAT12.000.01–100MiniMaxH3SigmaShift shift_video.
shift_audiooptFLOAT3.000.01–100MiniMaxH3SigmaShift shift_audio.
bd_grp_perfoptBDGROUP性能—
clear_vram_between_segmentsoptBOOLEANtrue段间清理显存:每段结束后卸载模型并清空 CUDA 缓存。
export_source_imagesoptBOOLEANfalse输出 source_images(时间轴原片帧对比)。默认关以节省内存。
preview_taeoptBOOLEANtrue真彩实时预览:采样中用 taeh3.safetensors 小解码器输出真彩动画预览帧。需要 models/vae_approx/taeh3.safetensors;缺失或加载失败时自动回退 latent2rgb 快速预览。
preview_overhead_percentoptFLOAT101–50动画预览开销上限(%):采样中动画预览的解码+编码耗时不超过生成时间的该百分比。默认 10%;设得越高预览越流畅,但整体出片耗时越多。
sag_enabledoptBOOLEANtrue加速开关:开=Comfy Kitchen INT8 加速注意力(速度/显存接近 SageAttention,且与 H3 兼容);关=PyTorch 注意力(精度最稳,速度/显存开销更大)。注意:SageAttention 的 int8 路径与 H3 长序列不兼容(灰噪/崩坏),已不再使用。
modeloptMODEL可选:外挂 UNETLoader。连接后优先于内置「模型加载」配置。
video_vaeoptVAE可选:外挂 Video VAE。连接后优先于内置配置。
audio_vaeoptVAE可选:外挂 Audio VAE。连接后优先于内置配置。
clipoptCLIP可选:外挂 CLIP。连接后优先于内置配置。
model_refoptMODEL可选:外挂 ref2va UNET。连接后优先于内置配置。
ltx_modeloptMODEL可选:外挂 LTX 2.5 精修模型。连接后优先于内置配置。
ltx_clipoptCLIP可选:外挂 LTX CLIP。连接后优先于内置配置。
ltx_vaeoptVAE可选:外挂 LTX VAE。连接后优先于内置配置。
ltx_upscaleroptLATENT_UPSCALE_MODEL可选:外挂 LTX latent 放大器。连接后优先于内置配置。
ltx_iclora_downscaleoptFLOAT11–8可选:外挂 IC-LoRA 加载器的 latent_downscale_factor。连接后优先于内置配置。
ltx_iclora_modeloptMODEL可选:外挂 IC-LoRA 补丁模型。连接后优先于内置配置。

Outputs (5)

NameTypeDescription
imagesIMAGE—
audioAUDIO—
fpsFLOAT—
compare_p1IMAGE—
compare_p2IMAGE—