MiniMax H3提示词扩写指令 V2·完整模式(Tlant)
Every H3 switch in one node, nothing left to chance
- 图像
- LLM扩写指令
- 图像
Here's the whole V2 design in one sentence: stop letting the node guess. The V1 family (Simple Prompt, the six config nodes, the Assembler) fills anything you leave alone with seed-stable random values. MiniMax H3提示词扩写指令 V2·完整模式(Tlant) throws all of that out. It's a single node holding every category - identity, face, editing, camera, scene, audio - where every single option is explicit, nothing is random, and unset means unset, not "roll the dice."
The mechanism is a clean three-valued system applied to every optional element:
不指定- the default. The element is simply not written into the instruction. The LLM is explicitly told an absent option is genuinely unspecified and must not invent a value to fill a template.无- explicit prohibition. The element is banned; for unavoidable visual properties (like a face existing), it means "don't deliberately change the source."自行推断- the LLM is asked to infer a safe, source-grounded value when relevant.
V2 has no local seed at all - no random roll happens inside the node, so reproducibility comes from repeating your settings, not from a lucky draw. What you don't select stays out of the prompt. It's the "explicitness over convenience" philosophy, and for a model as prompt-sensitive as H3 it's a defensible trade.
Common inputs for both V2 nodes: 视频时长 (4–15s), 主题预设 + 自定义主题 (non-empty custom theme overrides the preset completely), and 提示词长度最小值 - default 200, the minimum English words the final H3 prompt must contain. The node enforces it via the instruction, not by padding: the LLM is told to reach the minimum through useful detail and never by repetition.
The full list is big - fifty-odd dropdowns covering identity lock, facial/head motion, face direction, gaze, expression, cuts, shot count and scale, transitions, continuity, ending, every official camera move, amplitude/speed/energy, handheld, axis, depth of field, motion blur, background, scene expansion, lighting, atmosphere, hair, clothing, VFX, environment events, and the whole audio stack (mode, music, tempo, energy, beat sync, SFX, ambience, dialogue, language, voice style, lip sync).
Two audio policies are worth calling out because they're the design's best feature. Music is opt-in: the LLM may only generate background music if you explicitly choose music or a full mix, enable music, pick concrete music parameters, or set them to infer - otherwise the instruction forces non_diegetic_music: N/A. Dialogue is off by default: no invented speech, singing, narration, or on-screen text unless you ask. Both stop the "why did it add a soundtrack/voice I never wanted" complaints that plague generative video.
Outputs: LLM扩写指令 (STRING) and 图像 (IMAGE passthrough) - feed both into a vision LLM (Tlant's Llama Server Chat is the natural bridge) to get the clean final H3 prompt.
Install
ComfyUI Manager → search ComfyUI-Tlant-Toolkit → install, or:
cd ComfyUI/custom_nodes
git clone https://github.com/Tlant/ComfyUI-Tlant-Toolkit
Restart ComfyUI. No dependencies, no model downloads.
Gotchas
An image or a non-empty 图像描述 is required, same as V1 - no image, no description, and the node refuses. And don't fight the philosophy: if you leave 身份锁定 on 不指定, the prompt won't mention identity strength at all, and H3's baseline continuity still applies but you've given up your say. That's not a bug, it's the point. Also remember V2 coexists with V1 in the same pack - they're separate nodes, and V2 won't change how your old V1 workflows behave.
Inputs (52)
| Name | Type | Default | Description |
|---|---|---|---|
| 视频时长 | INT | 104–15 | MiniMax H3 官方支持 4–15 秒。 |
| 主题预设 | COMBO | 不指定 | 视频的整体创意方向;自定义主题非空时会完全覆盖此项。 |
| 自定义主题 | STRING | 自由填写视频主题;只要非空就完全覆盖主题预设,不与预设拼接。 | |
| 提示词长度最小值 | INT | 2001–2000 | 要求远程 LLM 输出的完整英文 MiniMax H3 提示词不少于多少 words。 |
| 身份锁定 | COMBO | 不指定 | 控制人物身份和五官稳定强度;无不会取消 I2VA 最基本的一致性。 |
| 面部动作 | COMBO | 不指定 | 控制面部动作幅度;无表示不生成面部动作。 |
| 头部动作 | COMBO | 不指定 | 控制转头和头部动作幅度。 |
| 面部朝向 | COMBO | 不指定 | 控制人物面部相对镜头的朝向。 |
| 视线行为 | COMBO | 不指定 | 控制眼睛注视方向和是否跟随镜头。 |
| 表情 | COMBO | 不指定 | 控制表情保持、眨眼、微笑或情绪变化。 |
| 切镜模式 | COMBO | 不指定 | 控制是否允许或必须使用多个镜头。 |
| 镜头数量 | COMBO | 不指定 | 指定总镜头数量;无按单镜头解释。 |
| 景别 | COMBO | 不指定 | 指定特写、中近景、中景、全景或混合景别。 |
| 景别变化模式 | COMBO | 不指定 | 控制不同镜头之间的景别变化顺序。 |
| 切镜节奏 | COMBO | 不指定 | 控制切镜在时间线上的疏密或是否跟随节拍。 |
| 转场方式 | COMBO | 不指定 | 控制硬切、匹配剪辑、甩镜、遮挡、溶解或淡入淡出。 |
| 连续性 | COMBO | 不指定 | 控制严格连续、轻微跳时或广告/MV蒙太奇。 |
| 结尾方式 | COMBO | 不指定 | 控制视频最后一刻如何收束。 |
| 运镜类型 | COMBO | 不指定 | 使用中文选择 MiniMax H3 官方支持的运镜类型。 |
| 运镜幅度 | COMBO | 不指定 | 控制构图变化范围。 |
| 运镜速度 | COMBO | 不指定 | 控制镜头移动速度。 |
| 镜头能量 | COMBO | 不指定 | 控制镜头整体稳定或动态程度。 |
| 手持感 | COMBO | 不指定 | 控制手持抖动强度;无表示稳定机位。 |
| 镜头轴线 | COMBO | 不指定 | 控制正面轴线、小角度变化或环绕。 |
| 景深 | COMBO | 不指定 | 控制浅、中、深景深或保持原图。 |
| 运动模糊 | COMBO | 不指定 | 控制动态画面中的运动模糊。 |
| 背景锁定 | COMBO | 不指定 | 控制背景保持、扩展、局部变化或重构。 |
| 场景扩展 | COMBO | 不指定 | 控制是否允许生成原图画框外的场景信息。 |
| 背景运动 | COMBO | 不指定 | 控制背景元素的动态幅度。 |
| 光线变化 | COMBO | 不指定 | 控制光线保持、扫光、明暗、闪烁、颜色或昼夜变化。 |
| 氛围效果 | COMBO | 不指定 | 控制风雨雪雾、烟尘、粒子、散景或光晕。 |
| 头发运动 | COMBO | 不指定 | 控制头发是否随动作或风产生运动。 |
| 服装运动 | COMBO | 不指定 | 控制衣物的自然、风吹或动作驱动变化。 |
| 视觉特效 | COMBO | 不指定 | 控制视觉特效强度;无表示不添加特效。 |
| 环境事件 | COMBO | 不指定 | 选择场景中可见的环境事件。 |
| 音频模式 | COMBO | 不指定 | 控制静音、环境声、动作声、音乐或完整混音。 |
| 音乐启用 | COMBO | 不指定 | 显式关闭时优先于其他音乐设置;默认不指定也不会自动生成配乐。 |
| 音乐风格 | COMBO | 不指定 | 具体选择音乐风格会被视为明确启用背景音乐。 |
| 音乐速度 | COMBO | 不指定 | 控制非叙事背景音乐速度。 |
| 音乐能量 | COMBO | 不指定 | 控制音乐强度及其随时间的变化。 |
| 节拍同步 | COMBO | 不指定 | 控制人物动作、切镜或两者是否跟随音乐节拍。 |
| 音效丰富度 | COMBO | 不指定 | 控制动作音和场景音效数量,必须有画面依据。 |
| 环境声预设 | COMBO | 不指定 | 不指定时不写入配置;无表示禁止环境声;自行推断只允许画面有依据的环境声。 |
| 自定义环境声 | STRING | 非空时覆盖环境声预设;填写希望出现的具体环境声。 | |
| 对话预设 | COMBO | 不指定 | 默认不生成对白;自行推断时也只允许在主题确有必要时生成一句很短的对白。 |
| 自定义对话 | STRING | 非空时覆盖对话预设;原文会被要求逐字保留,不自动翻译或改写。 | |
| 声音风格预设 | COMBO | 不指定 | 仅在实际存在对白、演唱或旁白时生效。 |
| 自定义声音风格 | STRING | 非空时覆盖声音风格预设,例如:低沉、平静、略带沙哑、语速缓慢。 | |
| 声音语言 | COMBO | 不指定 | 仅在存在对白、演唱或旁白时规定语言。 |
| 口型同步 | COMBO | 不指定 | 控制人物发声时是否生成对应口型。 |
| 图像opt | IMAGE | 连接后,实际图片是远程视觉模型的最高优先级依据,并会原样输出。 | |
| 图像描述opt | STRING | 无图时必须提供;有图时只作为辅助描述,冲突时以图片为准。 |
Outputs (2)
| Name | Type | Description |
|---|---|---|
| LLM扩写指令 | STRING | — |
| 图像 | IMAGE | — |