H3 Context-IR 编译器
The H3 Node That Won't Let an LLM Number Your References
- compiled_prompt
- compile_report
Every prompt-enhancer workflow eventually produces the same mess: the model returns lovely English with the labels invented. <Picture 3> for a picture you renumbered, an <Audio 4> that doesn't exist, a [Shot 2] cut point at a timecode that doesn't match your second shot. H3 is order-sensitive about all of it, and it doesn't warn you - it just renders something that ignores half your references.
H3_ContextCompiler is the node that takes that job away from the model. The LLM supplies English semantics only. The program writes the sections, the reference labels, the Subject and Speaker IDs, the shot cut points, and the dialogue verbatim from your locked draft - which is exactly the split our LLM-in-ComfyUI piece describes as the pattern that actually holds up: narrow the model's job, then chain it rather than trusting one free-form rewrite.
How it works
Two required inputs. context_json is the locked context object the workbench emitted; llm_response is the raw text your LLM returned. That's the whole surface.
On execution it does json.loads on the context and demands schema_version be h3-context-2. A context from the deprecated 1.0.0 prototype, or one you hand-edited, stops there with context_json 不是 h3-context-2。 In Ref2VA it also re-derives the presentation map from the draft's asset order and refuses to continue if the stored map doesn't match - so you can't compile against numbering that belongs to a different wiring.
Then it parses the enrichment response against a closed vocabulary. Allowed keys are schema_version, style_en, shots (each with shot_no, visual_en, camera_en, diegetic_sound_en, dialogue_cues_en), overall_soundscape_en, non_diegetic_music_en, subject_descriptions_en and asset_notes_en. Anything else is rejected outright. English is enforced on every semantic field, asset_notes_en must cover every slot the draft declares, and invented slot names bounce. A stray markdown code fence around the JSON is tolerated, so pasting straight out of a chat window usually works.
Then the deterministic renderer builds the prompt. Base modes get an optional alignment instruction first - I2VA pins <Picture 1> to the 0.00-second mark, L2VA pins it to the end, FL2VA states both ends with the real effective duration - followed by integrated_multimodal_description:, then overall_soundscape: and non_diegetic_music:. Shots read [Shot 1] and then [Shot 2] At 00:04.000,; dialogue is emitted as ... (S1) says with restrained irritation: <d>[Chinese] 你来了。</d>, language-tagged by script and copied character-for-character. Ref2VA switches to the six-section form: subject_definitions: → summary: → retention_analysis: → detailed_description: → overall_soundscape: → non_diegetic_music:.
Inputs and outputs
context_json- fromH3_ContextWorkbench, thecontext_jsonoutput. Multiline string, so pasting works too.llm_response- the enrichment JSON as text. Wire it from an LLM node if you have one, or just paste it into the widget after running the prompt in a browser tab.compiled_prompt- the finished H3 prompt. Goes toH3_PromptAuditand then to your H3 generation node's prompt input.compile_report- one line confirming the mode and stating that the deterministic renderer wrote the numbering, timeline and dialogue.
The llm_role and llm_prompt outputs on the workbench are your system and user messages. Any provider that returns the requested JSON will do - the pack isn't tied to RunningHub, so local Qwen or a GGUF model is fine. If you're using a chat model, this is the reason to prefer small-and-obedient over a reasoning model: the job is format-following, and chain-of-thought scratch-work is exactly the kind of thing that leaks into a field that must be plain English.
Install
ComfyUI Manager, search H3 Context Compiler. Or:
cd ComfyUI/custom_nodes
git clone https://github.com/wanski24hours-cmyk/h3-prompt-compiler.git
Restart. The pack ships no pip dependencies and downloads no models - the H3 checkpoint itself is a separate concern. The author's own dev checks are pytest -q and python -m compileall -q . from the pack root, if you want to sanity-check the install.
Where people get burned
Never hand-edit compiled_prompt between this node and the audit. It's tempting to nudge one number; the audit compares against the draft inside context_json and will fail on the token you touched, plus whatever you broke downstream. Fix the draft and re-run the chain instead.
Re-run the workbench when your assets change. A context_json cached from before you added a picture carries the old presentation map, so Ref2VA compilation stops, and base-mode numbering may silently refer to a slot that moved. And if the LLM's output is rejected, read the error - it names the field. Usually it's a Chinese character in a *_en field, an asset note missing for a slot you declared, or a <Picture N> tag the model wrote when it was told not to.
Last thing: the compiler only trusts locked dialogue. The model never writes the actual line, only the delivery cue, so if a line looks wrong in the output, it's wrong in your draft.
Inputs (2)
| Name | Type | Default | Description |
|---|---|---|---|
| context_json | STRING | {} | — |
| llm_response | STRING | — |
Outputs (2)
| Name | Type | Description |
|---|---|---|
| compiled_prompt | STRING | — |
| compile_report | STRING | — |