VRGDG LongShot Auto Director Context
VRGDG LongShot Auto Director Context
- planner_instructions
What it is
Despite the name, this node does not call an LLM. There's no API key, no model download, no network access. It builds one very long, very opinionated prompt - a string - that asks whatever LLM you do wire it to for a complete story plan for a MiniMax H3 LongShot: characters with permanent IDs, per-chunk action, camera moves, dialogue, and the handoff state at the end of each chunk.
That's the whole idea, and it's a good one. Chunked video generation fails in a boring way: each clip gets prompted in isolation, the character's jacket changes colour, the geography teleports, and the last chunk ends somewhere the first one never set up. Planning the story once and forcing every later prompt to inherit it is the fix.
It sits in the LongShot chain - Auto Director Context → an LLM node → Plan Extractor (one per chunk) → LLM Prompt Context → your H3 sampler. Separately, Keyframe Director Context turns the same plan into four coordinated stills.
How it works
Everything is templated into the request text, and the constraints are the interesting part - they're numbers your downstream chunks get built around.
- Duration math.
total_chunks × chunk_durationbecomes the story's total runtime, printed into the prompt so the LLM budgets beats to it. - Dialogue budget. Words-per-second rates are baked in - none 0.0, light 1.1, normal 1.8, heavy 2.2 - and multiplied by chunk duration, floored. "Heavy" at 7.5 seconds gives you about 16 newly spoken words per chunk, total across all speakers. The prompt says contractions count as one word and to leave room to breathe. This is the single most useful guardrail in here; unconstrained LLMs write 60 words of dialogue into a 7.5-second clip every single time.
- Story mode briefs. Each of the 12 modes carries its own instruction -
movie trailerwants escalation and a hook with no title cards,advertisementforbids unsupported claims and visible text,surreal art filmdemands dream logic that still preserves continuity. - Editing and audio contracts. With
allow_cutsoff, every chunk must be one uninterrupted physical take; on, cuts are allowed but continuity isn't negotiable.built-in audiomeans H3 generates synced sound, so the LLM should invent exact dialogue now.custom audiomeans a soundtrack exists already - plan the words, but they'll have to be recorded or synthesised into that track before video generation. - Speaker IDs. The plan is required to assign S1, S2 … and the schema in the prompt demands a
charactersarray withid,name,visual_description,voice_description. Everything downstream keys off those IDs, so don't let a model skip them.
Inputs worth setting
story_mode and user_idea are the two that decide what you get. Leave user_idea blank and the node tells the model to invent the concept completely - fine for a test, useless for an actual project.
Then total_chunks (1–16, default 4) and chunk_duration (1–60s, default 7.5). If you're working to a real soundtrack, make those two multiply out to your audio length, because every later node assumes the number you set here. dialogue_amount sets the word budget above.
reference_setup, image_1_role and image_2_role describe the up-to-two planning images you'll attach on the LLM node itself. The node writes a strict contract: analyse only visible details, never claim to see what isn't there, and if you need multiple characters plus a location, prefer one multi-character cast sheet and one location shot rather than two solo portraits.
creative_seed picks a repeatable creative direction - a dice roll you can re-roll deterministically, not a diffusion seed. tone, genre and visual_style are the rest of the dial.
One output: planner_instructions, a STRING. Wire it to the pack's VRGDG LLM Multi node and attach your reference images there.
Install
Manager → search vrgamedev, or:
cd ComfyUI/custom_nodes
git clone https://github.com/vrgamegirl19/comfyui-vrgamedevgirl.git
python -m pip install -r comfyui-vrgamedevgirl/requirements.txt
Restart and hard-refresh. Fair warning on the requirements: this pack installs llama-cpp-python and voxcpm, both of which can want a real compiler toolchain. If pip starts compiling, install Cython and scikit-build-core first (the README says exactly this), and prefer Python 3.12 over 3.13 on a Windows portable build.
When it goes wrong
The node itself has one realistic failure mode: it raises a ValueError naming the bad widget if you pass an enum value it doesn't know, which in practice means a stale workflow from an older pack version.
Everything else is the LLM's problem. The plan must come back as valid JSON with no Markdown fences, and a small local model will happily wrap it in prose, apologise, or emit trailing commentary. The downstream extractors are fence-tolerant and slice between the first { and last }, but chatty output around a second JSON object will break them. Use a model that follows format instructions, turn on JSON mode if your provider has it, and treat a failed parse as a model problem, not a node problem.
And be realistic about the ceiling: a 3B local model here gets you a coherent set of chunks, not a story you'd want to watch. Plan first, generate second - and let the planning step be where the stronger model earns its cost.
Inputs (14)
| Name | Type | Default | Description |
|---|---|---|---|
| story_mode | COMBO | random cinematic story | 12 options: random cinematic story, user-directed story, reference-inspired story, movie trailer, advertisement, cartoon short, +6 |
| user_idea | STRING | — | |
| genre | COMBO | automatic | 10 options: automatic, drama, thriller, horror, science fiction, fantasy, +4 |
| visual_style | COMBO | automatic | 7 options: automatic, live action, 2D cartoon, 3D animation, anime, stop motion, +1 |
| tone | STRING | cinematic, emotionally engaging | — |
| dialogue_amount | COMBO | normal | 4 options: none, light, normal, heavy |
| audio_mode | COMBO | built-in audio | 2 options: built-in audio, custom audio |
| allow_cuts | BOOLEAN | false | — |
| total_chunks | INT | 41–16 | — |
| chunk_duration | FLOAT | 7.51–60 | — |
| reference_setup | COMBO | no reference images | 3 options: no reference images, image 1 only, images 1 and 2 |
| image_1_role | COMBO | protagonist | 7 options: protagonist, second character, multi-character cast sheet, location, product, creature or prop, +1 |
| image_2_role | COMBO | second character | 7 options: protagonist, second character, multi-character cast sheet, location, product, creature or prop, +1 |
| creative_seed | INT | 10–2147483647 | — |
Outputs (1)
| Name | Type | Description |
|---|---|---|
| planner_instructions | STRING | — |