Nodes/Vision Prompt Assistant/Vision Prompt — Visual Director
ComfyUI Node

Vision Prompt — Visual Director

One Prompt for a Whole Storyboard Sheet — and It Won't Draw a Single Panel

By elgalardi·Created 2 months ago·Updated a day ago· 3
Vision Prompt — Visual Director
  • llm
  • reference_image_1
  • reference_image_2
  • reference_image_3
  • reference_image_4
  • reference_video
  • prompt
  • sheet_plan
  • validation
  • usage_stats
◄request►
◄sheet_typeStoryboard►
◄scene_count9►
◄layoutgrid►
◄aspect_ratio16:9►
◄annotationspanels_only►
◄annotation_languageEnglish►
◄seed0►
◄max_tokens4096►
◄temperature0.70►
◄image_max_dimension1536►
◄video_samples5►
◄hold_promptfalse►
◄target_modelMiniMax H3►
◄bypassfalse►
◄direction_context—►

The name says "Visual Director", which invites you to expect an image. It isn't one. It's an LLM that reads your rough brief - plus up to four reference images, and optionally a sampled video - and writes a single prompt describing an entire sheet: a storyboard grid, a character turnaround, a custom layout, or an edit instruction. Nothing is generated, no tensors pass through, your canvas is not resized.

Why you'd want it

Sheets beat one-prompt-per-image for a mechanical reason: describe every panel in one prompt, render it in one pass, and the cast, lighting and palette share a latent context instead of drifting across runs - which is why turnaround sheets have survived two complete tool turnovers in this ecosystem (character-consistency.md).

Writing that by hand is tedious: same faces re-described identically in nine places, camera language kept consistent, gutters the model actually respects. That's narrow structured-rewriting work, exactly what a language model is for (llm-in-comfyui.md).

How it works

Wire an LLMMODEL provider in. The pack ships three: H3OpenRouterModel (key field or OPENROUTER_API_KEY), H3OllamaModel for a local server, and H3LLMModelAPI for any OpenAI-compatible chat endpoint. The director serialises your widgets into a JSON brief, attaches reference images as data URLs downscaled to image_max_dimension, and makes exactly one chat request with a strict JSON schema, reasoning disabled, and your max_tokens / temperature / seed.

Then it renders the model's JSON back into text: for a grid it computes the layout itself (columns = ceil(sqrt(panel count)), plus horizontal and vertical variants) and emits one block per panel with a PANEL n: description and VIEW: camera line, alongside a shared visual_bible. A truncated response (finish_reason: length) or a wrong panel count errors out rather than handing you half a sheet - raise max_tokens or drop panels.

The controls you'll actually touch

request is the brief. Empty works only if references are connected.

sheet_type picks the mode: Storyboard, Character Sheet, Custom Sheet, Edit or Image. Edit is the odd one: it needs reference_image_1 as the source plus a written request, ignores every panel and layout control, and preserves untargeted content by instruction, not by a mask - so connect that source to your downstream editor too.

scene_count (labelled "Scenes / Panels", 1–24) is the panel count in that one image, and layout is grid / horizontal / vertical. aspect_ratio describes the whole sheet, not individual panels, and it isn't wired to anything - set your image workflow's resolution to match by hand.

annotations is the one people get wrong in both directions: panels_only means no lettering at all, brief_labels gives short captions, production_notes gives ACTION / CAMERA / AUDIO notes. annotation_language affects visible lettering only; production prose stays English.

target_model is grammar, not provider: MiniMax H3 keeps <Picture N> / <Subject N> identifiers, Qwen writes one prose paragraph using <imageN> and puts sizing advice in the plan metadata instead of the prompt. Don't switch it with Hold on - held prompts from the other target are rejected.

reference_video takes a decoded IMAGE batch (VHS output, not a file path) and samples 3, 5 or 10 frames uniformly - experimental, no audio, and those frames are evidence only, never generation references.

Outputs

prompt goes into your text-conditioning path. sheet_plan is the plan the LLM returned, as JSON - visual bible, subject definitions, retention analysis, panels, warnings. validation is human-readable status plus warnings. usage_stats is token counts and the model name.

Two switches matter. hold_prompt reuses this node's last saved prompt and plan without calling the LLM; every input change is silently ignored while it's on, which is the most confusing thing about the node. The hold lives in output/Sexy AI Studio/director_hold/plans.json, keyed to the node's id. bypass passes request through to prompt unchanged, with no LLM call and an empty sheet_plan, and it beats Hold. And seed seeds the LLM request, not your sampler.

Install

cd ComfyUI/custom_nodes
git clone https://github.com/elgalardi/ComfyUI-VisionPromptAssistant

Or search Vision Prompt Assistant in ComfyUI Manager and restart. Zero pip dependencies, no model downloads for this node, but it requires ComfyUI 0.30.0+ (requires-comfyui in the pack's pyproject.toml) because it's written against the modern comfy_api.latest node API - an older build won't load it.

Where it goes wrong

Planning a sheet and not wiring the references to the generator. It writes text. Identity retention depends on the image model actually seeing your reference images, in the same order the prompt describes them.

Hold left on. You changed a control, got identical output, filed a bug.

Sequence awareness. The README flags this as future work: the director plans everything in one call and never inspects the generated segments, so continuity across chained clips is promised, not verified.

Cost and credentials. External providers bill per call. Keys belong in the loader or the environment, never in a workflow you share - the standard caution for anything holding a credential and phoning out by design (external-api-nodes.md).

Categorytext/vision_prompt

Inputs (22)

NameTypeDefaultDescription
requestSTRING—
sheet_typeCOMBOStoryboard4 options: Storyboard, Character Sheet, Custom Sheet, Edit
scene_countINT91–24—
layoutCOMBOgrid3 options: grid, horizontal, vertical
aspect_ratioCOMBO16:95 options: 16:9, 4:3, 1:1, 3:4, 9:16
annotationsCOMBOpanels_only3 options: panels_only, brief_labels, production_notes
annotation_languageSTRINGEnglish—
seedINT00–4294967295—
max_tokensINT4096512–16384—
temperatureFLOAT0.700–1—
image_max_dimensionINT1536512–2048—
video_samplesCOMBO53 options: 3, 5, 10
hold_promptBOOLEANfalseReuse this node's last saved sheet without calling the LLM. Disable to apply any changes.
target_modelCOMBOMiniMax H3Selects prompt grammar, not the LLM provider or image model loader.
bypassBOOLEANfalsePass request through unchanged. Skip LLM, Hold, references and all planning controls.
llmoptLLMMODEL—
reference_image_1optIMAGE—
reference_image_2optIMAGE—
reference_image_3optIMAGE—
reference_image_4optIMAGE—
reference_videooptIMAGEExperimental: decoded video IMAGE batch, e.g. from VHS. Uniform frame samples only; no audio analysis.
direction_contextoptSTRINGConnect H3 Compact Direction Controls. Motion becomes static visual staging; audio becomes optional annotations.

Outputs (4)

NameTypeDescription
promptSTRING—
sheet_planSTRING—
validationSTRING—
usage_statsSTRING—