沐阳 H3 · 导演台(全功能)
One node that runs a whole storyboarded H3 shoot
- h3
- model
- sampler
- media
- ref_video
- ref_audio
- 前段视频
- 前段音频
- 二采模型
- 一采成片
- 一采成片音频
- images
- audio
- plan_json
- fps
This is the pack's flagship, the node the whole thing is named after. 沐阳 H3 · 导演台(全功能) (H3Director) is a single node that organizes the entire job - storyboard, media, generation, seams, and optionally a second-pass upscale - into one runnable unit. The bundled quick-start workflow is only seven nodes: 沐阳 H3 加载器, 沐阳 H3 取模型, a KSampler select, this, and a video output. Everything else the README talks about as separate machinery - the Media Agent, the segment plan, the anchor long video, Turbo scheduling, second pass - is orchestrated from inside it.
Underneath it's not magic: the Director reuses the same existing Myang nodes as building blocks, so old workflows and the individual advanced nodes don't get replaced. But instead of hand-wiring a dozen nodes per segment, you fill in storyboards or hand the Director a script, and it builds the timeline.
Two ways to build a timeline
- Manual storyboard (
source_modedefaults to导演台分镜卡) - each shot card gets a title, prompt, duration, and its own media. Durations snap to H3's17k+5frame grid. One card accepts up to 9 images, 3 videos, and 3 audio files. There's a shared-media area for character images, locations, and music used throughout; media on a card belongs to that shot. Refer to items as@图片1,@视频1,@音频1. - Agent / long-script split - with
llm_enabled, an LLM builds the timeline from total duration, per-segment duration, the media inventory, and writing rules.script_fallbackholds the script (or convert that widget to an input and connect an Agentmyang_prompt).skill_preset(defaultauto) picks a writing skill for the segment-prompt structure, andskill_textis your own rules on top. With the LLM disabled, the input prompt takes the local splitting path and burns zero tokens.
task_mode picks the generation job: pure generation (no reference video), motion transfer (each segment follows its slice of the reference video - a motion-transfer card can pick its own action source, else the global ref_video is used), or video continuation. llm_service and vlm_service control which configured LLM/VLM backs the split and whether a VLM first describes the shared media so the LLM can decide what to reference.
The settings that matter most
- h3 / model / sampler - the loader bundle, the first-pass model (base or Turbo-joined both go here), and a sampler node.
- resolution / aspect_ratio / steps / denoise / scheduler / noise_seed - the generation contract;
simplescheduler and denoise 1.0 are the Turbo-friendly defaults. - context_length - the continuity window (22 default; same scale as everywhere else in the pack).
- 起始段 (start_segment) - resume support for motion transfer:
1re-runs everything; on a resume you set the first un-generated segment and connect the previous segment's finished clip to前段视频, whose tail only seeds the anchor (it never enters the ref2va reference channel).前段音频is its optional soundtrack for the audio seam. - The whole 二采… family - second-pass mode, resolution, steps, denoise, upscaler, chunking, passes and seed strategy, all inline.
二采模型takes the Ref2VA base model before the Turbo LoRA. - save_segments / segment_prefix / save_raw_segments - per-segment output so a crash doesn't cost you everything.
Outputs: images, audio, the plan_json it executed, and fps.
LLM services
Open ComfyUI Settings → Myang_node → LLM Service Settings. It supports OpenAI-compatible APIs and Ollama, multiple URL/API-key routes per service, round-robin or primary-route-first selection, and cooldown/failover after rate limits and timeouts. Keys are never returned to the browser, and config lives in user/default/Myang_node/config/llm_services.json.
Install
ComfyUI Manager search "ComfyUI-MiniMaxH3-Myang", or git clone https://github.com/civilcoco/ComfyUI-MiniMaxH3-Myang into custom_nodes, restart, hard-refresh. No extra Python deps. Load example_workflows/Minimax_H3_Myang_Director_CN.json, pick your own models, media, prompts and seed, render two segments at a 22-frame context, inspect the seam - picture and sound - before scaling up. And the standing H3 license reminder: local weights are geofenced out of the US, EU, UK and South Korea.
Inputs (91)
| Name | Type | Default | Description |
|---|---|---|---|
| h3 | MYANG_H3 | — | |
| model | MODEL | 一采模型;基础模型或 Turbo 联合模型都直接接这里 | |
| sampler | SAMPLER | — | |
| source_mode | COMBO | 导演台分镜卡(手动逐镜头) | 2 options: 导演台分镜卡(手动逐镜头), Agent / 长剧本智能切分 |
| timeline_json | STRING | {"version": 2, "shots": [{"id": "shot_1", "enabled": true, "duration_seconds": 5.0, "brief": "镜头 1", "prompt": "", "transition": "开场", "asset_mode": "仅本镜头", "assets": []}]} | — |
| script_fallback | STRING | 智能切分时作为剧本;可把此控件转换为输入后连接 Agent easy_prompt | |
| total_seconds | FLOAT | 601–3600 | — |
| segment_seconds | FLOAT | 10.04–20 | — |
| llm_enabled | BOOLEAN | true | — |
| llm_service | COMBO | 1 options: 未配置 LLM 服务 | |
| task_mode | COMBO | 纯生成(不用参考视频) | 3 options: 动作迁移(跟随参考视频), 视频续写(接着往下演), 纯生成(不用参考视频) |
| resolution | COMBO | 480P | 12 options: 360P, 416P, 480P, 540P, 640P, 720P, +6 |
| aspect_ratio | COMBO | 16:9 | 8 options: 1:1, 2:3, 3:2, 3:4, 4:3, 9:16, +2 |
| width | INT | 86432–16384 | — |
| height | INT | 48032–16384 | — |
| steps | INT | 251–200 | — |
| denoise | FLOAT | 1.000.01–1 | — |
| scheduler | COMBO | simple | 3 options: simple, beta, normal |
| noise_seed | INT | 00–18446744073709550000 | — |
| context_length | COMBO | 22 | 4 options: 22, 5, 39, 56 |
| ref_image_size | COMBO | 匹配生成分辨率 | 5 options: 匹配生成分辨率, 最大1K面积, 最大1.5K面积, 最大2K面积, 匹配素材(原尺寸) |
| 二采开启 | BOOLEAN | false | — |
| 二采模式 | COMBO | 放大 + 二采(推荐) | 3 options: 放大 + 二采(推荐), 同分辨率二采(不放大), 仅放大(不二采·最快) |
| 二采分辨率 | COMBO | 832P | 9 options: 540P, 640P, 720P, 768P, 832P, 928P, +3 |
| 二采自定义宽 | INT | 166432–8192 | — |
| 二采自定义高 | INT | 92832–8192 | — |
| 二采步数 | INT | 41–100 | — |
| 二采重绘幅度 | FLOAT | 0.200.01–1 | — |
| 二采调度器 | COMBO | beta | 3 options: beta, simple, normal |
| 二采采样器 | COMBO | res_multistep | 2 options: res_multistep, euler |
| 二采放大方式 | COMBO | neural_3d (神经3D Latent放大·推荐) | 3 options: neural_3d (神经3D Latent放大·推荐), pixel (像素放大·自用版工作流方式), nvidia_rtx_vsr (NVIDIA RTX 视频超分·实验) |
| 二采分块帧数 | INT | 41–64 | — |
| 二采Latent模型 | COMBO | 神经3D放大权重;其他放大方式会自动隐藏 | |
| 二采精度 | COMBO | fp16(推荐·省显存) | 3 options: fp16(推荐·省显存), fp32(最高稳定性), bf16(实验) |
| 二采时间分块 | INT | 00–256 | 神经3D放大的时间分块;0=全上下文单次推理(无接缝,推荐)。显存不够再往上调,8 最省显存 |
| 二采轮数 | INT | 11–8 | — |
| 二采种子策略 | COMBO | 每轮沿用同一种子 | 2 options: 每轮沿用同一种子, 每轮种子 +1 |
| save_segments | BOOLEAN | true | — |
| segment_prefix | STRING | video/H3_导演台 | — |
| save_raw_segments | BOOLEAN | false | — |
| 参考视频分辨率 | COMBO | 匹配参考视频原分辨率 | 动作迁移/视频续写的参考视频预处理;默认保持原尺寸,最高 1080P |
| 参考视频自定义宽 | INT | 192032–1920 | — |
| 参考视频自定义高 | INT | 108032–1920 | — |
| 起始段 | INT | 11–12 | 勾选『从指定段开始』后生效;支持动作迁移和导演台手动分镜卡。Agent 智能切分始终从第 1 段生成 |
| skill_preset | COMBO | auto | 智能切分的写作技能:决定每段提示词的输出结构、分镜格式和素材标签写法。auto 先用一次很短的调用按剧本选技能 |
| skill_text | STRING | 自定义写作规则,排在所选技能之前,优先级最高 | |
| vlm_service | COMBO | off | 开启后先让 VLM 看一遍每个公共素材,把画面内容写进清单,LLM 才能按内容判断每段该引用哪个素材 |
| 脸部精修开启 | BOOLEAN | false | — |
| 脸部检测器 | COMBO | bbox\face_yolov8m.pt | 1 options: bbox\face_yolov8m.pt |
| 脸部精修步数 | INT | 41–50 | — |
| 脸部精修重绘 | FLOAT | 0.450.01–1 | — |
| 脸部裁剪倍率 | FLOAT | 2.51.2–8 | — |
| 脸部身份图序号 | INT | 00–9 | 0=自动跟踪最大脸;1~9=用对应 @图片N 辅助锁定人物 |
| 动作修复开启 | BOOLEAN | false | — |
| 动作修复档位 | COMBO | balanced (default) | 3 options: balanced (default), max quality (wide plateau), economy (tight spans) |
| 动作修复步数 | INT | 64–50 | — |
| 动作修复注入 | FLOAT | 0.700.05–1 | — |
| 多视角分镜开启 | BOOLEAN | false | — |
| 多视角角色图片序号 | INT | 11–9 | — |
| 多视角尺寸 | COMBO | 512 | 2 options: 512, 1024 |
| 多视角步数 | INT | 284–50 | — |
| 多视角LoRA | COMBO | minimax_h3_five_view_1024cont_s600.safetensors | 1 options: minimax_h3_five_view_1024cont_s600.safetensors |
| 多视角LoRA强度 | FLOAT | 0.750–2 | — |
| 二采复用一采条件 | BOOLEAN | true | 开:只复用一采的文本 token 与已编码参考素材,不包含、不复制640P一采成片;二采目标latent始终独立。还省掉每段一次文本编码和参考素材VAE编码。关:按二采分辨率重跑条件,参考图会被重采样到更大面积,token 全变,低降噪几步收不过去,容易涂抹和轻微身份漂移 |
| 音频精修开启 | BOOLEAN | false | 冻结视频 latent,仅用未挂 Turbo LoRA 的基模附加去噪音频。每一步仍接近一次完整 H3 前向计算 |
| 音频精修步数 | INT | 41–100 | — |
| 音频去噪强度 | FLOAT | 0.500.01–1 | 音频重新加噪后再修复的深度。数值越大,改动越明显;不是画面重绘,也不会开放视频 latent |
| 音频精修采样器 | COMBO | euler | 2 options: euler, res_multistep |
| 音频精修调度器 | COMBO | simple | 3 options: simple, beta, normal |
| 音频接缝平滑 | BOOLEAN | true | 用下一段被裁掉的重叠锚点音频融合上一段尾部;不重复声音,不改变成片时长 |
| 音频接缝时长 | FLOAT | 800–500 | 建议 60~120ms;对白密集可缩短,环境声可适当加长 |
| 从指定段开始 | BOOLEAN | false | 关闭时无条件从第 1 段开始;开启后才读取『起始段』 |
| 二采显存策略 | COMBO | 自动平衡(16GB推荐) | 16GB建议自动平衡;DynamicVRAM下数值表示留给二采激活的容量,并非永久空置显存;避免Windows共享显存换页,不降低832P最终输出质量 |
| 二采自定义显存预留 | FLOAT | 1.250–8 | 仅自定义档生效,单位GB;DynamicVRAM下是激活空间,非动态加载器下是传统显存预留;0=沿用ComfyUI启动设置 |
| 二采自定义预览间隔 | INT | 20–100 | 仅自定义档生效;0=关闭二采逐步清晰预览,1=每步,2=每2步;最终预览始终保留 |
| 二采连续Sigma | BOOLEAN | false | 把一采和二采步数组成一条连续噪声轨迹;不增加总步数。实验模式复用一采模型/采样器/调度器,只支持 neural_3d 放大或同分辨率二采。 |
| mediaopt | MINIMAX_H3_MEDIA | — | |
| ref_videoopt | IMAGE | — | |
| ref_audioopt | AUDIO | — | |
| 前段视频opt | IMAGE | 断点续跑时上一段已生成的成片。若起始分镜为『承接』,只用其结尾做上下文锚点,不会作为 ref2va 参考视频;起始分镜为『切镜』时不需要连接 | |
| 前段音频opt | AUDIO | 可选:上一段成片的音轨,用于声音接缝 | |
| 二采模型opt | MODEL | 导演台二采使用的 Ref2VA 基模;接 Turbo LoRA 之前的模型 | |
| 粗剪时间轴开启opt | BOOLEAN | false | 开启后按粗剪工程的时长模式生成;可沿用导演台时长自动补齐另一端,也可由时间轴 I/O 覆盖本次生成时长;成片精确覆盖该区间 |
| 粗剪工程opt | STRING | {"format":"myang.roughcut","version":1,"id":"roughcut_4b68647cd38b4ea1aa554eec74c67484","revision":0,"settings":{"fps":24.0,"width":1920,"height":1080,"auto_align":true,"snap_enabled":true,"writeback_enabled":false},"selection":{"in_frame":0,"out_frame":120,"in_set":false,"out_set":false,"anchor_side":"","duration_mode":"director","target_track":"video_1","start_mode":"off","end_mode":"off","use_first_frame":false,"use_last_frame":false},"tracks":[{"id":"video_1","kind":"video","name":"V1","clips":[]},{"id":"audio_1","kind":"audio","name":"A1","clips":[]}]} | 粗剪工作区的结构化工程数据;请用导演台里的『打开时间轴』编辑 |
| 动作迁移自动分段opt | BOOLEAN | true | 仅动作迁移生效;关闭后自动匹配动作视频长度,只生成一段 |
| 一采显存策略opt | COMBO | 自动平衡(16GB推荐) | 16GB推荐自动平衡:条件编码完成后卸载CLIP/VAE,采样时不逐步调用视频VAE;每段采样完成仍显示清晰预览。兼容模式保留原来的每步清晰预览。 |
| 一采断点模式opt | COMBO | 关闭 | 保存模式会按分段前缀保存完整音画latent;读取模式按原段号直接进入二采。成片兼容入口只支持单段。 |
| 一采成片opt | IMAGE | 兼容入口:加载一段已保存的一采视频,内部VAE重编码后直接二采 | |
| 一采成片音频opt | AUDIO | 可选:与一采成片配套的原始音频 | |
| 二采后VSR增强opt | BOOLEAN | false | 仅『同分辨率二采』生效:解码后按原尺寸跑一遍 NVIDIA RTX VSR,只去噪锐化,不改分辨率。需要 nvvfx 与 NVIDIA 显卡;每段一次隔离进程加一整段磁盘往返、逐帧推理。『放大 + 二采』和『仅放大』请用放大方式里的 VSR |
| 分层提示词opt | BOOLEAN | false | 仅『Agent / 长剧本智能切分』生效:每段的声音与台词各自单独生成,篇幅短、更少报错、也更好逐层审查;台词按本段真实秒数换算成硬性字数上限,放不下的顺延到下一段,不会再把一整段对话塞进一个 8 秒镜头 |
Outputs (4)
| Name | Type | Description |
|---|---|---|
| images | IMAGE | — |
| audio | AUDIO | — |
| plan_json | STRING | — |
| fps | FLOAT | — |