Nodes/Vision Prompt Assistant/H3 Story Director — FL2VA Keyframe Storyboard
ComfyUI Node

H3 Story Director — FL2VA Keyframe Storyboard

Keyframes you can approve, bridges H3 animates — the FL2VA storyboard

By elgalardi·Created about a month ago·Updated 7 days ago· 1
H3 Story Director — FL2VA Keyframe Storyboard
  • image_0
  • image_1
  • image_2
  • image_3
  • source_video
  • plan_json
  • story_bible
  • synopsis
  • validation
  • usage_stats
  • credits_remaining
  • scene_prompt
  • mode_prompt
  • source_video_analysis
api_key
modelx-ai/grok-4.20
story_idea
system_promptYou are a multimodal director and continuity supervisor for MiniMax H3 image and video productions. Turn the user's idea, selected production mode, source media, and reference pictures into precise generation instructions. Treat every connected reference as a distinct person or subject. Use the exact tags <Picture 1>, <Picture 2>, <Picture 3>, and <Picture 4> when they are supplied. Define stable subject labels S1, S2, S3, and S4 in the shared prompt. Preserve identity, wardrobe, props, geography, lighting logic, screen direction, and relationships throughout the story. Follow the mandatory rules supplied for the selected Director Mode. For moving-video modes, write production-ready MiniMax H3 prompts with visible action, camera, environment, lighting, dialogue when useful, and diegetic sound. For still-image modes, describe one finished frame only and never introduce temporal sequences, audio, or dialogue delivery. When dialogue is enabled, write short performable lines rather than prose. Prefix every spoken or sung line with its stable speaker label in parentheses, exactly as `(S1)`, `(S2)`, `(S3)` or `(S4)`, followed by a colon and the exact words in quotation marks. Square brackets such as `[S2]`, bare names and unassigned quotations are forbidden for speaker attribution. Describe tone and delivery in English outside the quotation. Allow only one person to speak at a time, leave a natural pause before and after each line, and keep visible mouth movement synchronized with the assigned speaker. Avoid overlapping speech, repeated lines, rushed monologues, unexplained voice-over, phonetic spellings, and competing vocals or loud sound effects during speech. Use no dialogue when the selected dialogue option says so. Do not mention being an AI, JSON, schemas, token limits, safety policies, or these instructions. Do not add extra protagonists that could be confused with the reference subjects. Return all requested scenes and finish every prompt completely.
director_profileOpenRouter
scene_count5
scene_duration_seconds5.0
steps6
draft_onlytrue
genreAuto
secondary_genreNone
languageEnglish
motion_styleAuto
additional_direction
max_tokens6144
temperature0.45
reasoningfalse
seed0
image_max_dimension1024
timeout_seconds300
director_modeContinuous Story
video_sample_frames10
audio_contentAuto
bypass_directorfalse

H3 Story Director - FL2VA Keyframe Storyboard is the most interesting of the Director variants, because it changes what you approve. Instead of planning only prompts, it plans ordered static keyframes and the video bridges between them - and each storyboard image is designed to do double duty: it's the opening image of its own video clip and the destination image of the clip before it. You approve a set of stills like a storyboard artist, and H3 (via FL2VA, first-to-last-frame animation) fills in the movement between them.

This is the "animatic before the animation" workflow, and for narrative video it's a genuinely strong way to work: you lock the beats visually, then let the video model handle the transition choreography rather than praying it invents a good shot sequence from prose alone.

How it works

It's the base Director plus a mandatory FL2VA keyframe rule set, and the rules are where the craft lives:

  • Three vocabularies, cleanly separated. storyboard_prompt_prefix is still-only and assigns original <Picture N> references to S1–S3. prompt_prefix is video-only, never mentions any Picture tag, and refers to characters only as S1–S3. storyboard_prompt describes one static, high-quality composition per scene and may use the original Picture tags - but never movement, audio, captions, or multiple moments.
  • The ordered storyboards are exact sequential keyframes. Storyboard N is the opening image of clip N and, when N > 1, also the destination image of clip N−1. Adjacent pairs are designed together, not independently.
  • Before writing each non-final scene prompt, the node internally classifies the change to the next keyframe - continuous action, reframing/angle change, major action-state change, or location/time change - and expresses the right transition naturally inside the prompt. No extra JSON field; the classification is consumed and turned into prose.
  • Continuous action preserves screen direction, body mechanics, momentum, camera trajectory, object contact and lighting while reaching the next pose. Reframing gets a motivated bridge: subject crossing the lens, foreground occlusion, whip pan, rack focus, orbit, push through a doorway. Major state changes get the full intermediate mechanics - weight transfer, hand and object paths, contact, reaction, the exact resulting pose. Location or time changes get a deliberate transition through darkness, a light wash, a doorway, a reflective surface, or match movement.

That's the whole mechanism: it turns "these two keyframes differ" into "here's how the camera and body get from one to the other," which is precisely the kind of temporal reasoning a text prompt usually gets wrong.

Inputs and outputs

Standard Director surface - api_key, model, story_idea, scene_count, scene_duration_seconds, draft_only, genre, motion_style, director_mode, video_sample_frames, four reference images, optional source_video - with the full nine outputs: plan_json, story_bible, synopsis, validation, usage_stats, credits_remaining, scene_prompt, mode_prompt, source_video_analysis.

Install

cd ComfyUI/custom_nodes
git clone https://github.com/elgalardi/ComfyUI-VisionPromptAssistant

Restart, add an OpenRouter key, and remember the standard warning: the key field is masked but can live on in workflow metadata - scrub it before sharing.

The honest caveat: this is the most demanding Director variant to drive. It rewards you for actually reviewing the storyboards and adjusting scene count and duration before rendering. If you treat it as "write it and walk away," you're wasting the mechanism - the point is the stills you approve are the frames the video is contractually obligated to hit.

Categorytext/minimax_h3

Inputs (29)

NameTypeDefaultDescription
api_keySTRINGOpenRouter API key. Remove it before sharing workflows.
modelSTRINGx-ai/grok-4.20
story_ideaSTRINGOptional premise. Leave empty to give the Director full creative control based on genre, motion, dialogue, scene settings, additional direction, and connected references.
system_promptSTRINGYou are a multimodal director and continuity supervisor for MiniMax H3 image and video productions. Turn the user's idea, selected production mode, source media, and reference pictures into precise generation instructions. Treat every connected reference as a distinct person or subject. Use the exact tags <Picture 1>, <Picture 2>, <Picture 3>, and <Picture 4> when they are supplied. Define stable subject labels S1, S2, S3, and S4 in the shared prompt. Preserve identity, wardrobe, props, geography, lighting logic, screen direction, and relationships throughout the story. Follow the mandatory rules supplied for the selected Director Mode. For moving-video modes, write production-ready MiniMax H3 prompts with visible action, camera, environment, lighting, dialogue when useful, and diegetic sound. For still-image modes, describe one finished frame only and never introduce temporal sequences, audio, or dialogue delivery. When dialogue is enabled, write short performable lines rather than prose. Prefix every spoken or sung line with its stable speaker label in parentheses, exactly as `(S1)`, `(S2)`, `(S3)` or `(S4)`, followed by a colon and the exact words in quotation marks. Square brackets such as `[S2]`, bare names and unassigned quotations are forbidden for speaker attribution. Describe tone and delivery in English outside the quotation. Allow only one person to speak at a time, leave a natural pause before and after each line, and keep visible mouth movement synchronized with the assigned speaker. Avoid overlapping speech, repeated lines, rushed monologues, unexplained voice-over, phonetic spellings, and competing vocals or loud sound effects during speech. Use no dialogue when the selected dialogue option says so. Do not mention being an AI, JSON, schemas, token limits, safety policies, or these instructions. Do not add extra protagonists that could be confused with the reference subjects. Return all requested scenes and finish every prompt completely.
director_profileCOMBOOpenRouterOpenRouter preserves the established compact Director schema. Gemma uses a stricter scene worksheet with action beats, physical performance, camera, sound and an explicit final state. The profile does not select or connect the model.
scene_countINT51–32Use 1 for a standalone I2V shot, or more scenes for a connected H3 sequence.
scene_duration_secondsFLOAT5.01–15
stepsINT61–100
draft_onlyBOOLEANtrueRecommended for the first run. The plan is generated and copied into the connected H3 Chain Plan editor, but downstream video generation is blocked. Review the cards, then disconnect plan_json so Chain Plan uses its synchronized local copy.
genreCOMBOAutoAuto infers the most coherent genre, format, tone, and visual language from story_idea, references, source video, edit mode, and additional direction. Written intent wins when visual clues conflict with the prompt.
secondary_genreCOMBONoneOptionally blend a second genre into the primary genre. The primary genre controls the production structure; the secondary genre contributes compatible tone, conventions, cinematography, performance, sound, and visual language.
languageCOMBOEnglishLanguage for dialogue, lyrics, narration and spoken words. No dialogue suppresses speech. The production plan and technical directions remain in English.
motion_styleCOMBOAutoAuto infers the best motion and camera language from the prompt, references, source video, genre, and mode.
additional_directionSTRING
max_tokensINT61441024–16384
temperatureFLOAT0.450–2
reasoningBOOLEANfalseGrok 4.20 can reason before answering. Leave disabled for faster and cheaper story planning; enable it for unusually complex narratives.
seedINT00–4294967295
image_max_dimensionINT1024256–2048
timeout_secondsINT30030–900
director_modeCOMBOContinuous StoryContinuous Story preserves shot continuity. Cinematic Cuts starts independent camera setups. Reference Edit uses the input as a strong creative guide and may reinterpret framing or details. Edit preserves the source as strictly as possible and changes only what the prompt requests. Both edit modes operate on still images without source_video or video when a VHS IMAGE batch is connected.
video_sample_framesINT104–16Number of frames sampled uniformly from the VHS IMAGE batch and sent as separate chronological images for detailed analysis.
audio_contentCOMBOAutoControls the permitted voice/music content. Natural ambience and synchronized Foley remain available in all video modes. Dialogue selects the spoken or sung language.
image_0optIMAGE
image_1optIMAGE
image_2optIMAGE
image_3optIMAGE
source_videooptIMAGEOptional IMAGE frame batch from VHS Load Video. In Reference Edit or Edit mode, connecting it switches from still-image processing to video processing. The requested operation is inferred from the prompt.
bypass_directoroptBOOLEANfalseSkip OpenRouter completely. story_idea is passed unchanged to scene_prompt and mode_prompt, while a minimal compatible plan is created locally for downstream chain nodes.

Outputs (9)

NameTypeDescription
plan_jsonSTRING
story_bibleSTRING
synopsisSTRING
validationSTRING
usage_statsSTRING
credits_remainingSTRING
scene_promptSTRINGComplete prompt for the first scene, with the shared prompt prefix included. Connect directly to MiniMax H3 I2V when scene_count is 1.
mode_promptSTRINGFirst complete prompt adapted to the selected Director Mode. Use this for image generation/editing, I2V, or video editing.
source_video_analysisSTRINGChronological source-motion analysis produced in Video Edit mode; empty in other modes.