Nodes/VRGameDevGirl Video Enhancement Nodes/VRGDG LongShot Auto Director Context
ComfyUI Node

VRGDG LongShot Auto Director Context

VRGDG LongShot Auto Director Context

By vrgamegirl19·Created about a year ago·Updated a day ago· 742
VRGDG LongShot Auto Director Context
    • planner_instructions
    ◄story_moderandom cinematic story►
    ◄user_idea►
    ◄genreautomatic►
    ◄visual_styleautomatic►
    ◄tonecinematic, emotionally engaging►
    ◄dialogue_amountnormal►
    ◄audio_modebuilt-in audio►
    ◄allow_cutsfalse►
    ◄total_chunks4►
    ◄chunk_duration7.5►
    ◄reference_setupno reference images►
    ◄image_1_roleprotagonist►
    ◄image_2_rolesecond character►
    ◄creative_seed1►

    What it is

    Despite the name, this node does not call an LLM. There's no API key, no model download, no network access. It builds one very long, very opinionated prompt - a string - that asks whatever LLM you do wire it to for a complete story plan for a MiniMax H3 LongShot: characters with permanent IDs, per-chunk action, camera moves, dialogue, and the handoff state at the end of each chunk.

    That's the whole idea, and it's a good one. Chunked video generation fails in a boring way: each clip gets prompted in isolation, the character's jacket changes colour, the geography teleports, and the last chunk ends somewhere the first one never set up. Planning the story once and forcing every later prompt to inherit it is the fix.

    It sits in the LongShot chain - Auto Director Context → an LLM node → Plan Extractor (one per chunk) → LLM Prompt Context → your H3 sampler. Separately, Keyframe Director Context turns the same plan into four coordinated stills.

    How it works

    Everything is templated into the request text, and the constraints are the interesting part - they're numbers your downstream chunks get built around.

    • Duration math. total_chunks × chunk_duration becomes the story's total runtime, printed into the prompt so the LLM budgets beats to it.
    • Dialogue budget. Words-per-second rates are baked in - none 0.0, light 1.1, normal 1.8, heavy 2.2 - and multiplied by chunk duration, floored. "Heavy" at 7.5 seconds gives you about 16 newly spoken words per chunk, total across all speakers. The prompt says contractions count as one word and to leave room to breathe. This is the single most useful guardrail in here; unconstrained LLMs write 60 words of dialogue into a 7.5-second clip every single time.
    • Story mode briefs. Each of the 12 modes carries its own instruction - movie trailer wants escalation and a hook with no title cards, advertisement forbids unsupported claims and visible text, surreal art film demands dream logic that still preserves continuity.
    • Editing and audio contracts. With allow_cuts off, every chunk must be one uninterrupted physical take; on, cuts are allowed but continuity isn't negotiable. built-in audio means H3 generates synced sound, so the LLM should invent exact dialogue now. custom audio means a soundtrack exists already - plan the words, but they'll have to be recorded or synthesised into that track before video generation.
    • Speaker IDs. The plan is required to assign S1, S2 … and the schema in the prompt demands a characters array with id, name, visual_description, voice_description. Everything downstream keys off those IDs, so don't let a model skip them.

    Inputs worth setting

    story_mode and user_idea are the two that decide what you get. Leave user_idea blank and the node tells the model to invent the concept completely - fine for a test, useless for an actual project.

    Then total_chunks (1–16, default 4) and chunk_duration (1–60s, default 7.5). If you're working to a real soundtrack, make those two multiply out to your audio length, because every later node assumes the number you set here. dialogue_amount sets the word budget above.

    reference_setup, image_1_role and image_2_role describe the up-to-two planning images you'll attach on the LLM node itself. The node writes a strict contract: analyse only visible details, never claim to see what isn't there, and if you need multiple characters plus a location, prefer one multi-character cast sheet and one location shot rather than two solo portraits.

    creative_seed picks a repeatable creative direction - a dice roll you can re-roll deterministically, not a diffusion seed. tone, genre and visual_style are the rest of the dial.

    One output: planner_instructions, a STRING. Wire it to the pack's VRGDG LLM Multi node and attach your reference images there.

    Install

    Manager → search vrgamedev, or:

    cd ComfyUI/custom_nodes
    git clone https://github.com/vrgamegirl19/comfyui-vrgamedevgirl.git
    python -m pip install -r comfyui-vrgamedevgirl/requirements.txt
    

    Restart and hard-refresh. Fair warning on the requirements: this pack installs llama-cpp-python and voxcpm, both of which can want a real compiler toolchain. If pip starts compiling, install Cython and scikit-build-core first (the README says exactly this), and prefer Python 3.12 over 3.13 on a Windows portable build.

    When it goes wrong

    The node itself has one realistic failure mode: it raises a ValueError naming the bad widget if you pass an enum value it doesn't know, which in practice means a stale workflow from an older pack version.

    Everything else is the LLM's problem. The plan must come back as valid JSON with no Markdown fences, and a small local model will happily wrap it in prose, apologise, or emit trailing commentary. The downstream extractors are fence-tolerant and slice between the first { and last }, but chatty output around a second JSON object will break them. Use a model that follows format instructions, turn on JSON mode if your provider has it, and treat a failed parse as a model problem, not a node problem.

    And be realistic about the ceiling: a 3B local model here gets you a coherent set of chunks, not a story you'd want to watch. Plan first, generate second - and let the planning step be where the stronger model earns its cost.

    CategoryVRGDG/Video/Long Shot

    Inputs (14)

    NameTypeDefaultDescription
    story_modeCOMBOrandom cinematic story12 options: random cinematic story, user-directed story, reference-inspired story, movie trailer, advertisement, cartoon short, +6
    user_ideaSTRING—
    genreCOMBOautomatic10 options: automatic, drama, thriller, horror, science fiction, fantasy, +4
    visual_styleCOMBOautomatic7 options: automatic, live action, 2D cartoon, 3D animation, anime, stop motion, +1
    toneSTRINGcinematic, emotionally engaging—
    dialogue_amountCOMBOnormal4 options: none, light, normal, heavy
    audio_modeCOMBObuilt-in audio2 options: built-in audio, custom audio
    allow_cutsBOOLEANfalse—
    total_chunksINT41–16—
    chunk_durationFLOAT7.51–60—
    reference_setupCOMBOno reference images3 options: no reference images, image 1 only, images 1 and 2
    image_1_roleCOMBOprotagonist7 options: protagonist, second character, multi-character cast sheet, location, product, creature or prop, +1
    image_2_roleCOMBOsecond character7 options: protagonist, second character, multi-character cast sheet, location, product, creature or prop, +1
    creative_seedINT10–2147483647—

    Outputs (1)

    NameTypeDescription
    planner_instructionsSTRING—