Nodes/VRGameDevGirl Video Enhancement Nodes/VRGDG LongShot LLM Prompt Context
ComfyUI Node

VRGDG LongShot LLM Prompt Context

VRGDG LongShot LLM Prompt Context

By vrgamegirl19·Created about a year ago·Updated a day ago· 742
VRGDG LongShot LLM Prompt Context
    • llm_instructions
    ◄video_idea►
    ◄chunk_direction►
    ◄lyrics_this_chunk►
    ◄previous_lyric_carryover►
    ◄chunk_number1►
    ◄total_chunks4►
    ◄chunk_duration7.5►
    ◄context_frames22►
    ◄frame_setupprevious ending + last frame►
    ◄vocal_modesinging►
    ◄allow_cutsfalse►
    ◄audio_modecustom audio►
    ◄visual_reference_modenormal references►
    ◄storyboard_panel_count4►
    ◄storyboard_layoutautomatic►
    ◄storyboard_image_slotImage 3►
    ◄previous_chunk_prompt—►

    Where the LongShot chain gets real

    This is the heaviest node in the LongShot set, and the one you'll live in. Everything before it planned; this one assembles the instruction that makes an LLM write the finished MiniMax H3 prompt for one chunk - and it's the node that knows about frame counts, overlap padding and what the previous chunk ended up doing.

    The mental model: you loop this node once per chunk. Chunk 1 in, wire the resulting H3 prompt back into the next instance's previous_chunk_prompt, bump chunk_number, repeat.

    How it works

    The interesting mechanism is the timing math, and it's all in frames at a hard-coded 24 fps:

    • body = round(chunk_duration × 24) - the frames you actually want to keep.
    • prefix = context_frames (1, 5, 22 or 39) whenever the frame setup reuses the previous chunk's tail; otherwise 0.
    • requested = prefix + body, then generated = requested + (5 - requested) % 17 - the request is padded up to the next multiple of 17 frames, which is H3's chunking arithmetic.
    • endpoint = (prefix + body - 1) / 24 - the timestamp of the last frame you care about.
    • padding = generated - prefix - body - trailing frames you'll throw away.

    All of that is printed into the instruction, with an explicit instruction to place no required action in the discarded padding. That's the bit that stops a model from staging a beat at 7.4 s in an 7.5 s chunk when only 7.29 s survives the overlap trim.

    Two more pieces of continuity plumbing. previous_chunk_prompt (optional input, so you drag a link) is trimmed to the last ~1,800 characters of the previous prompt's action, with <d>...</d> dialogue replaced by [previous dialogue omitted] and old timing stripped. The prompt calls that "private continuity memory - never output": the model uses it for momentum, not to copy stale words.

    The rest of the instruction comes from a mode-selected cheat sheet. built-in audio tells the LLM to use H3's three-field integrated format (integrated_multimodal_description, overall_soundscape, non_diegetic_music). custom audio switches to the six-field format - subject_definitions, summary, retention_analysis, detailed_description, overall_soundscape, non_diegetic_music - with <Audio 1> marked fully_copy so the model reuses your soundtrack instead of regenerating it. Same split for frame_setup: each option maps to the matching H3 task (I2VA, FL2VA, continuation, continuation + endpoint), and allow_cuts flips between one take and timestamped [Shot N] blocks.

    One output: llm_instructions.

    Inputs worth setting

    video_idea and chunk_direction come straight from the Plan Extractor. lyrics_this_chunk is the exact words for this chunk - quoted verbatim - with previous_lyric_carryover for a line bleeding across the boundary. If you supply no words and vocal_mode isn't instrumental / no vocals, the node explicitly tells the LLM not to invent dialogue or lyrics, which is a nicer failure than a hallucinated chorus.

    chunk_number and total_chunks matter more than they look: the last chunk gets a different closing instruction ("resolve the action… no unrequested fade or cut"), every other chunk is told to end on something continuable. context_frames is your overlap budget - 22 is a shade under a second at 24 fps and the sensible default; 39 costs more context, 1 is nearly none. Wire llm_instructions into any instruction-following LLM (the pack's VRGDG LLM Multi works), and its text output goes to your H3 sampler.

    Then the frame/reference block: frame_setup decides which attached images are frame anchors, and if visual_reference_mode is storyboard grid you pick storyboard_image_slot, storyboard_panel_count (2–6) and storyboard_layout. Watch for the collision here - the node reserves Image 1 and Image 2 for anchors depending on frame_setup, and raises Image 2 is already assigned by frame_setup … if your storyboard slot lands on a reserved one. Same class of error if you pick a layout too small for your panel count (a 2x2 can't hold six).

    Install

    Manager → search vrgamedev, or:

    cd ComfyUI/custom_nodes
    git clone https://github.com/vrgamegirl19/comfyui-vrgamedevgirl.git
    python -m pip install -r comfyui-vrgamedevgirl/requirements.txt
    

    Restart and hard-refresh. Nothing here needs a model download, but the pack's requirements are heavy - llama-cpp-python and voxcpm in particular want a compiler if no wheel matches your Python. Cython and scikit-build-core first, Python 3.12 over 3.13 on Windows.

    When it goes wrong

    • chunk_number must be between 1 and total_chunks. You bumped the counter past the end of the story. It's a guard, not a crash.
    • Seams jitter or the action jumps between chunks. Almost always chunk_duration or context_frames differing between passes. Those numbers define the overlap contract; change them mid-loop and the panels stop lining up.
    • H3 repeats dialogue from the previous chunk. The carryover excerpt strips <d> blocks, but only if your LLM actually formatted dialogue that way. If you're hand-writing the previous prompt, keep the tags.
    • Chunk 1 gets no prefix. Expected - nothing precedes it. Only chunks 2+ take the context frames.

    This is a bespoke workflow for one model at one frame rate, and it makes no apology for it. If you're generating on H3 at 24 fps, it will save you a genuinely tedious afternoon of frame arithmetic. If you're on anything else, none of it applies.

    CategoryVRGDG/Video/Long Shot

    Inputs (17)

    NameTypeDefaultDescription
    video_ideaSTRING—
    chunk_directionSTRING—
    lyrics_this_chunkSTRING—
    previous_lyric_carryoverSTRING—
    chunk_numberINT11–999—
    total_chunksINT41–999—
    chunk_durationFLOAT7.50.1–60—
    context_framesCOMBO224 options: 1, 5, 22, 39
    frame_setupCOMBOprevious ending + last frame5 options: first frame only, first + last frames, previous ending only, previous ending + last frame, no frame anchors
    vocal_modeCOMBOsinging3 options: singing, spoken dialogue, instrumental / no vocals
    allow_cutsBOOLEANfalse—
    audio_modeCOMBOcustom audio2 options: custom audio, built-in audio
    visual_reference_modeCOMBOnormal references2 options: normal references, storyboard grid
    storyboard_panel_countINT42–6—
    storyboard_layoutCOMBOautomatic5 options: automatic, horizontal, 2x2, 2x3, 3x2
    storyboard_image_slotCOMBOImage 34 options: Image 1, Image 2, Image 3, Image 4
    previous_chunk_promptoptSTRING—

    Outputs (1)

    NameTypeDescription
    llm_instructionsSTRING—