VRGDG LongShot Keyframe Prompt Extractor
Turning One JSON Blob Into Four Usable Image Prompts
- keyframe_1_prompt
- keyframe_2_prompt
- keyframe_3_prompt
- keyframe_4_prompt
- continuity_bible
- status
The seam between your LLM and your image generator
The Keyframe Director gives you one JSON document containing four image prompts. Your image generator wants four separate strings going into four separate encode nodes. This node is the adapter, and it behaves like a strict parser: it validates before it hands anything downstream, and it fails loudly rather than shipping a truncated prompt into a 30-minute render.
That last part is the reason to use it instead of copy-pasting. An LLM that ignores the 1,200-character rule will otherwise give you a prompt your model silently truncates, and you find out when keyframe three looks wrong.
How it works
keyframe_plan goes in as a STRING. Then:
- Leading and trailing Markdown fences are stripped with a regex, case-insensitively.
- The text between the first
{and the last}is sliced out, which conveniently drops "Here's your JSON:" preambles and any trailing sign-off. json.loadsruns. On failure you get a real message:The Keyframe Director returned invalid JSON at line 4, column 12: .... Read it, it's usually a trailing comma or an unescaped quote inside a prompt.- Each keyframe object is indexed by its
keyframenumber, falling back to array position if the field is absent or unparseable. If any of 1–4 is missing, the node raisesThe Keyframe Director plan is missing keyframes: [3]. - Every
image_promptmust exist and be 1,200 characters or fewer, or you getKeyframe 3 image_prompt is 1341 characters; maximum is 1200.
Then it returns the six outputs. Four are the standalone prompts, in order. continuity_bible is the shared identity / wardrobe / location / lighting / camera-path text the Director was asked to write - you don't have to use it, but it's the cheapest way to keep separate generations honest: prepend it to each prompt, or keep it as a reference in a text node while you tweak. status is not decoration - it reports the validated character counts (Four keyframe prompts validated: K1=1104 chars, K2=987 chars, ...), which is the fastest sanity check that your LLM actually obeyed the brief instead of emitting four 400-character sketches.
Wiring it up
This node is deliberately generator-agnostic. The four prompt outputs go wherever you'd otherwise type a prompt: CLIPTextEncode positive inputs if you're generating the stills inside ComfyUI, an API or browser-based image node if that's your path, or four different generators if you want to compare. Nothing stops you running all four keyframes in parallel branches with different seeds.
The one thing it won't do is fix a bad plan. If the Director returned three prompts about a coffee shop and one about a desert, you get exactly that, cleanly parsed.
If you're generating inside ComfyUI, the natural companion is VRGDG LongShot Collect 4 Keyframes, which takes the four resulting images and re-exposes them under their LongShot roles plus a preview batch. Extractor on the front, collector on the back.
Install
Manager → search vrgamedev, or:
cd ComfyUI/custom_nodes
git clone https://github.com/vrgamegirl19/comfyui-vrgamedevgirl.git
python -m pip install -r comfyui-vrgamedevgirl/requirements.txt
Restart and hard-refresh the browser page. This node needs nothing beyond the pack itself - no models, no VHS - but you're installing the full dependency list to get it.
When it goes wrong
- "does not contain a JSON object." The model answered in prose, or returned a JSON array instead of an object. Re-ask; the Director prompt specifies an object with a
keyframesarray. - Trailing commentary breaks the parse. The brace-slicing handles a lot, but an LLM that writes "Here's the plan: {...} and here's a note: {...}" gives you two objects and an ambiguous slice. Ask for JSON-only output.
- Length errors. The usual cause is an over-verbose image model setting. Nothing here is configurable - the limit is a length check, not a truncation - so either shorten
extra_image_directionor use a model that follows instructions. - Ordering looks scrambled. The node trusts the
keyframefield over array position. If your model numbered them 4,3,2,1, you get them back in 1–4 order and the mapping stays correct, which is the desired behaviour but can be surprising when you're eyeballing raw JSON.
One caveat on the whole chain: the JSON discipline comes from the prompt, not from constrained decoding, so the quality of this step is entirely a function of the model you chose. A small local LLM will pass the parse most of the time and still not give you four frames of one continuous shot.
Inputs (1)
| Name | Type | Default | Description |
|---|---|---|---|
| keyframe_plan | STRING | — |
Outputs (6)
| Name | Type | Description |
|---|---|---|
| keyframe_1_prompt | STRING | — |
| keyframe_2_prompt | STRING | — |
| keyframe_3_prompt | STRING | — |
| keyframe_4_prompt | STRING | — |
| continuity_bible | STRING | — |
| status | STRING | — |