H3 通用导演工作台|自动识别模式
Write Your H3 Prompt Like a Director, Not a Prompt Engineer
- first_frame
- last_frame
- p1
- p2
- p3
- p4
- p5
- p6
- p7
- p8
- p9
- v1
- v2
- v3
- va1
- va2
- va3
- a1
- a2
- a3
- resolved_mode
- context_json
- llm_role
- llm_prompt
- upload_plan_cn
- validation_report
- target_duration_seconds
- h3_length_frames
- vision_sheet
- vision_sheet_labels
MiniMax H3 wants its prompt in a specific shape: named sections in a fixed order, media written as <Picture 1> / <Video 1> / <Audio 1>, cut points as [Shot 2] At 00:04.000,, dialogue wrapped in <d>. Everybody's first instinct is to ask Claude or Codex to "just write the prompt", and that's exactly where it falls apart - an LLM is good at describing a scene and terrible at counting, and H3's numbering is not what the numbers on your sockets say.
H3_ContextWorkbench is the front door that fixes that. You write the locked structure in Chinese - title, duration, asset list, shots, dialogue - and it decides the mode, checks your actual wiring, computes H3's real reference labels, and hands an LLM only the semantic half of the job.
What it actually does
Four things, all deterministic:
- Parses your director draft. A regex-based Chinese parser, not a model. Title, duration, the asset list, then each shot with its cast, used assets, scene, sound and dialogue.
- Resolves the mode.
AUTOreads the asset list and picksT2VA/I2VA/FL2VA/L2VA/Ref2VAfor you, or you pin one. - Reconciles declaration against reality. Every declared slot must be physically connected, and nothing may be connected that the draft doesn't declare - which is what stops you rendering with a stale reference left plugged in from the last run.
- Builds the payload. A locked context object, a system prompt and user prompt for your LLM, an upload plan, a labeled vision contact sheet, and the frame count H3 will really use.
That last one matters more than it sounds. H3 runs at 24 fps on a 17k+5 frame grid, so your nominal seconds get rounded up. An 8-second draft lands exactly on 192 frames; a 10-second draft becomes 243 frames, or 10.125 s of actual video. h3_length_frames is the number to feed your clip-length node - target_duration_seconds is just what you asked for.
The draft format
The director_draft_cn widget is a real format. The widget's default shows the keyframe shape:
标题:示例
总时长:8秒
【资产装填清单】
首帧:开场参考图
【镜头1】
时长:8秒
出场人物:角色甲
使用资产:首帧
画面内容:角色甲从静止状态抬眼,看向镜头外的人。镜头缓慢推近。
现场声音:衣料轻响和安静室内底噪。
对白:
角色甲:你来了。
非画内音乐:无
Ref2VA entries carry a role inline: 图片1:角色甲角色卡|人物|完全保留, 视频1原声:原视频环境声|音频复用|完整复用. Shot durations must sum to the total, total must be 4–15 seconds, and every speaker in 对白 must already be listed in 出场人物. Assets you declare but never list in a shot's 使用资产 are an error, not a warning - the draft has to be a plan, not a wishlist.
Inputs and outputs that matter
Three required: generation_mode (AUTO plus the five modes), contract_profile (currently one choice, MiniMax H3 Official 2026-08) and director_draft_cn.
Optional media is the confusing part. first_frame / last_frame are the keyframe pair, and p1–p9 are reference pictures. v1–v3 are IMAGE frame sequences, not a VIDEO type - that's what H3's ref_video_N input actually wants. va1–va3 are the synced soundtracks belonging to v1–v3 respectively (ref_video_audio_N), and a1–a3 are standalone audio (ref_audio_N). A VA without its matching V is a hard error, because H3's node splits video picture and video sound into two sockets, not because VA is a fourth kind of reference.
Wire context_json to the compiler node, llm_role and llm_prompt to whatever LLM you're using, and vision_sheet to a multimodal one - it's a contact sheet with tiles labeled P1, LAST_FRAME, V1@10% and so on, sampled at roughly 10%, 50% and 90% of each video sequence purely so the model can see what's in them. Those previews never become <Picture N>. Read upload_plan_cn yourself: it's the single loading truth, showing logical slot → physical input → final prompt tag in one place.
Install
No pip dependencies at all, and no model downloads for the pack itself. Either search H3 Context Compiler in ComfyUI Manager, or:
cd ComfyUI/custom_nodes
git clone https://github.com/wanski24hours-cmyk/h3-prompt-compiler.git
Restart, then search H3 Context Compiler in the node menu. You supply the LLM - RunningHub's chat, a local GGUF, Qwen, anything that returns the JSON the node asks for.
Where people get burned
The errors here are all "your plan and your graph disagree", and they're worth reading rather than clicking past, because the message names the exact slot. The common ones: mixing keyframe assets with Ref2VA references (AUTO refuses), a declared slot with nothing plugged in, a leftover connection your draft doesn't mention, a shot whose speaker isn't in the cast list, and shot durations that don't add up.
Ref2VA has the real fence around it: up to 9 pictures, 3 videos, 3 standalone audio, and no more than 12 independent files; each video 2–15 s with 15 s total; standalone audio 2–15 s with a 15 s total; audio can never be your only input; and only one video may be tagged as the edit source, which must be the first video in presentation order. Also, note that first_frame/last_frame don't exist in Ref2VA - keyframes go in picture slots with a role. Get that wrong and you'll argue with the node for ten minutes before reading the error properly. Everyone does it once.
Inputs (23)
| Name | Type | Default | Description |
|---|---|---|---|
| generation_mode | COMBO | AUTO | 6 options: AUTO, T2VA, I2VA, FL2VA, L2VA, Ref2VA |
| contract_profile | COMBO | MiniMax H3 Official 2026-08 | 1 options: MiniMax H3 Official 2026-08 |
| director_draft_cn | STRING | 标题:示例 总时长:8秒 【资产装填清单】 首帧:开场参考图 【镜头1】 时长:8秒 出场人物:角色甲 使用资产:首帧 画面内容:角色甲从静止状态抬眼,看向镜头外的人。镜头缓慢推近。 对白: 角色甲:你来了。 | — |
| first_frameopt | IMAGE | — | |
| last_frameopt | IMAGE | — | |
| p1opt | IMAGE | — | |
| p2opt | IMAGE | — | |
| p3opt | IMAGE | — | |
| p4opt | IMAGE | — | |
| p5opt | IMAGE | — | |
| p6opt | IMAGE | — | |
| p7opt | IMAGE | — | |
| p8opt | IMAGE | — | |
| p9opt | IMAGE | — | |
| v1opt | IMAGE | — | |
| v2opt | IMAGE | — | |
| v3opt | IMAGE | — | |
| va1opt | AUDIO | — | |
| va2opt | AUDIO | — | |
| va3opt | AUDIO | — | |
| a1opt | AUDIO | — | |
| a2opt | AUDIO | — | |
| a3opt | AUDIO | — |
Outputs (10)
| Name | Type | Description |
|---|---|---|
| resolved_mode | STRING | — |
| context_json | STRING | — |
| llm_role | STRING | — |
| llm_prompt | STRING | — |
| upload_plan_cn | STRING | — |
| validation_report | STRING | — |
| target_duration_seconds | FLOAT | — |
| h3_length_frames | INT | — |
| vision_sheet | IMAGE | — |
| vision_sheet_labels | STRING | — |