Nodes/Kinburg-Nodes/Phantas Storyboard 🎞
ComfyUI Node

Phantas Storyboard 🎞

Turn a sentence into a chain of consistent stills

By KinburgΒ·Created 3 months agoΒ·Updated 6 days agoΒ· 1
Phantas Storyboard 🎞
  • config
  • board
  • prompts
  • beats
  • durations
  • style
  • report
β—„briefβ–Ί
β—„count_modescenesβ–Ί
β—„count3β–Ί
β—„target_length0.0β–Ί
β—„durationsβ–Ί
β—„castβ–Ί
β—„style_notesβ–Ί
β—„prompts_overrideβ–Ί
β—„preferred_length5.2β–Ί
β—„write_beatstrueβ–Ί
β—„cachediskβ–Ί
β—„live_previewtrueβ–Ί
β—„unload_after_runconfig defaultβ–Ί
β—„system_styleYou are a director writing the STYLE BIBLE for a sequence of still keyframes that will be generated one at a time, by an image model, in separate calls that never see each other. These blocks are pasted into EVERY frame's prompt unchanged, so they may contain only what is true in every single frame. CRITICAL: never describe the sequence's story, its beginning, its ending, or the stages of any change. A frame prompt that mentions the whole arc makes the image model try to show the arc in one picture. Answer with EXACTLY these three labelled blocks, in this order, and nothing else: [STYLE]: one paragraph, look and craft only β€” genre or reference, lens and focal length, depth of field, lighting, colour grade, grain and texture, atmosphere. No story, no camera moves, no shot list. [CAST]: every person who appears, ONE PER LINE, as `Name β€” a full physical description`. [SUBJECT]: one or two sentences of non-human INVARIANTS β€” the location, the time of day, the vehicle or the props. Never the story, never a start or an end state, and not the people (they are the cast). [NEGATIVE]: a comma-separated list of faults to avoid. Always include text, subtitles, logos and watermarks; add only faults β€” blur, artifacts, distorted anatomy, extra limbs, extra fingers, style breaks. NEVER list anything the sequence is supposed to DO: if the subject transforms, words like "morphing", "transformation" or "shape change" must not appear here. THE CAST BLOCK IS THE MOST IMPORTANT THING YOU WRITE. Each line is pasted verbatim into the prompt of every frame that person appears in, and the image model has no memory of the other frames: someone described in five words is a different human being in every picture. Write each of them the way a casting note does β€” apparent age, build and height, face shape and its distinguishing features, skin tone, hair colour and the exact cut, eye colour, facial hair, wardrobe from head to foot with colours and materials, and anything they always carry. Two or three sentences each, minimum. If people are described to you in the material you were given, use THOSE people and keep their given names, details and wording; invent nobody. If nobody appears, write `[CAST]: none`. Write in English, plainly, no markdown emphasis, no commentary.β–Ί
β—„system_planYou are a director breaking a brief into KEYFRAMES for a continuous video sequence. A keyframe is a frozen moment. Between two consecutive keyframes runs one shot, which the video model will generate as the movement from the first to the second. So N keyframes describe N-1 shots. THE SEQUENCE IS ONE CONTINUOUS TAKE. There are no cuts anywhere in it. The camera may travel, push in, pull back, crane, orbit or follow, but it never jumps: a change of framing between two keyframes is a camera MOVE that the shot between them performs. Never plan a montage. For each keyframe give: - "framing": the shot size and camera angle at that instant (e.g. "wide low-angle three-quarter", "medium tracking profile", "close-up over the shoulder"). Consecutive framings must be reachable by a camera move. - "present": the names of the cast members visible in that frame, exactly as the cast block spells them. An empty list if the frame shows nobody. Never name anyone who is not in the cast, and never leave someone out who is on screen β€” this list decides whose description gets attached to the picture. - "state": what is frozen on screen at that instant β€” position, pose, form, what the light is doing. A description of a STILL. Never write a change, never write "begins to", "starts to" or "is about to". For each transition (there is exactly one fewer than the keyframes) give: - "beat": two or three sentences, present tense, saying what visibly HAPPENS between those two keyframes and what the camera does. This is read alone by another writer who cannot see the other beats, so it must stand completely on its own and never say "then", "next", "finally" or "meanwhile". - "weight": an integer from 1 to 5 for how much visible change this transition carries. 1 is a held moment with a slow drift, 5 is the largest change in the sequence. Weights set how long each shot runs, so spend them honestly. Spend the whole brief across the sequence: the last keyframe lands on the brief's endpoint and no earlier one may get there first. If a transformation completes at keyframe 2 of 6, the plan is wrong. Answer with JSON only.β–Ί
β—„system_frameYou are writing the prompt for ONE still image: a single keyframe of a video sequence. You are given the style bible, this frame's framing and state, and β€” when there is one β€” the prompt of the frame immediately before it, so the two pictures can be of the same world. Write ONE paragraph describing what is in THIS frame, as a photograph of a frozen instant: - the subject, its exact pose and position in the frame, its form and its surfaces - the framing you were given: shot size, camera angle, lens behaviour - the environment and what the light is doing at this instant Rules: - **Everyone on screen is NAMED and described in full.** The image model never sees the other frames, so "the man from the previous shot" or "the singer" produces a different person every time. Give each person present their name and their face, hair, build and wardrobe again, in this frame, using the cast block's own words rather than a summary of them. - A still has no time in it. Never write a change, a movement in progress, "begins to", "starts to", "is about to", or anything that happens before or after this instant. - Never mention the sequence, the other frames, the shot, the story or its ending. - If a previous frame's prompt is given, keep everything the brief did not change: the same wardrobe, the same location, the same time of day, the same light, the same lens. - Plain descriptive English, one paragraph, no headings, no markdown, no commentary.β–Ί
β—„shot_lyricsβ€”β–Ί

Phantas Storyboard is where a video idea stops being a paragraph and becomes pictures. You give it a brief - "a woman walks a night city, rain on neon" - and a local LLM writes the keyframe prompts and the beats between them, with the kind of structure that keeps a chain of stills looking like one world instead of six different movies. The full pipeline it feeds is: Phantas Storyboard β†’ Phantas (renders the keyframes) β†’ Morpheus Storyboard (writes the video prompts) β†’ Morpheus (renders the video).

How it works

Three LLM calls, all streaming into a Kinburg Live Log if live_preview is on:

  1. A style bible - [STYLE], [CAST], [SUBJECT], [NEGATIVE] blocks written once and stamped into every frame byte-for-byte. That's how the look survives across separate image-model calls that never see each other.
  2. A plan - JSON constrained by a grammar built from the frame count, so the model can't emit the wrong number of keyframes or transitions. Each keyframe gets a framing (shot size, angle), present (who's on screen), and state (what's frozen at that instant); each transition gets a beat and a weight (1–5, how much change it carries - the weights set the shot lengths).
  3. One call per frame, shown the bible, its own framing and state, whoever is in it, and the previous frame's prompt - so consecutive pictures are of the same world.

The cast input is the whole game

An image model has no memory between calls: "the singer" or "the man from the previous shot" is a different human being in every picture. Identity has to be re-stated in full, in the same literal words, in every frame the person appears in. So cast takes one Name - full physical description line per person, at cover-art length - age, build, face, hair, eyes, wardrobe head to foot. Those lines replace whatever the bible wrote and are stamped verbatim, and the plan's present list decides whose lines land in which frame (pasting the whole cast into a shot of an empty room is how you grow a person who shouldn't be there). Leave cast empty and the bible writes one itself - usable, but its wording rather than yours; and leave it empty also when the subject is supposed to change, since a fixed description would contradict the transformation.

Two rules are baked into the system prompts: a frame is a frozen moment (never "he begins to transform" - the change lives in the beat), and the chain is one continuous take (a framing change has to be a camera move, because the shot performs it between two keyframes).

The inputs you'll set

config - a Local LLM Settings (GGUF) bundle. Text only: this node writes and never looks at pictures, so a light text model is the right one here and it leaves VRAM for the sampler. brief is the whole creative input. count_mode picks your unit - frames (N pictures = Nβˆ’1 shots), scenes (S shots = S+1 pictures), or duration (shot count from target_length). durations can be left empty to let the planner weight every transition by how much change it carries. prompts_override is the edit loop: run it, read the prompts output, fix the one frame that came out wrong, paste the whole thing back with --- separators. cache: disk keys the LLM answers causally so re-running the graph doesn't rewrite prompts and invalidate finished frames.

Outputs: board (into Phantas), prompts (editable), beats (in exactly the format Morpheus Storyboard takes - wire it there and the arc is planned once, not twice), durations, style, and report.

Installing it

Same pack: ComfyUI Manager β†’ "Kinburg-Nodes", or git clone https://github.com/Kinburg/Kinburg-Nodes into ComfyUI/custom_nodes, restart. You supply the GGUF LLM yourself; llama-cpp-python is installed automatically by install.py. Set unload_after_run to unload on a small card - the renderer next in line needs the room.

CategoryKinburg-Nodes/Bestiary/Phantas

Inputs (18)

NameTypeDefaultDescription
configKINBURG_LLM_CONFIGA 'Local LLM Settings (GGUF)' bundle. Text only β€” this node writes, it never looks at pictures, so a light text model is the right one here and leaves the VRAM for the sampler.
briefSTRINGWhat the clip is. One or several sentences: who or what is on screen, where, and what happens over the whole sequence from its first moment to its last.
count_modeCOMBOscenesWhich unit you are counting in. A keyframe sits BETWEEN shots, so the numbers are always one apart: β€’ frames β€” 'count' is how many KEYFRAMES to draw (N pictures = N-1 shots). β€’ scenes β€” 'count' is how many SHOTS to make (S shots = S+1 pictures). β€’ duration β€” neither; the shot count is worked out from 'target_length'.
countINT31–32Keyframes or scenes, depending on 'count_mode'. Ignored when the mode is 'duration'.
target_lengthFLOAT0.00–600Target length of the finished clip, in seconds. 0 = no target. In 'duration' mode this decides the shot count. In the other two it only shapes the shot LENGTHS, and it must be achievable: n shots can add up to between nΓ—5.17 s and nΓ—15.08 s, and a target outside that band is an error rather than something quietly clamped.
durationsSTRINGSeconds per shot: one value for all of them, or a comma list ('5.17, 8, 5.17') where the last value repeats. LEAVE EMPTY to let the planner decide β€” it weights every transition by how much change it carries, and the weights are laid onto H3's 0.71 s frame grid here.
castoptSTRINGWho is in this clip, ONE PER LINE, as 'Name β€” a full physical description'. Paste the character cards you already use for cover art: these lines are stamped VERBATIM into every frame that person appears in, and that is what makes the same face come out of every frame. Describe them at cover-art length β€” age, build, face, hair, eyes, wardrobe head to foot. A person described in five words is a different human being in each picture, because the image model never sees the other frames. Leave empty and the style-bible call writes the cast itself from the brief and from whatever the LLM Settings' context holds β€” usable, but its own wording rather than yours. Leave empty ALSO when the subject is supposed to change (a transformation), since a fixed description would contradict it.
style_notesoptSTRINGExtra instructions for the look only β€” reference films, lens, grade, era. Goes to the style-bible call, not to the plan.
prompts_overrideoptSTRINGPaste the 'prompts' output back here after editing it, and those frames are used verbatim instead of being written again. Frames are separated by a line of '---'; an empty entry means 'write this one'.
preferred_lengthoptFLOAT5.25.166666666666667–15.08333333333333How long an average shot should run. Sets the shot count in 'duration' mode, and the average length when no target is given.
write_beatsoptBOOLEANtrueWrite the beats β€” what visibly HAPPENS between each pair of keyframes β€” and the weights that give each shot its length. ON (the default) is the video path. The beats come out on the 'beats' output in exactly the format 'Morpheus Storyboard' takes, so the arc is planned once instead of twice. OFF is the slideshow path: use it when the keyframes are going to 'Save Clip' as slides and no video is being rendered. The planning grammar drops the transitions array altogether, so the model cannot spend tokens on directions nobody will read. Off also swaps the planner's own system prompt for the slideshow one, which is the half that changes what the pictures look like: with no video between them the CUT IS FREE, so the plan stops being one continuous take and starts being an edit β€” shot sizes jump, scale alternates, and some frames are given nobody at all. Type your own text into 'system_plan' and yours is used in both modes instead. The board still carries one shot per gap (blank beat, even length), so it stays a valid chain: turn this back on later, or wire the board to Morpheus anyway and its Storyboard writes the beats itself from the brief.
cacheoptCOMBOdiskCache the LLM's answers on disk, keyed causally, so re-running the graph does not rewrite the prompts and invalidate finished frames. Editing the brief re-rolls everything; editing one frame re-rolls that frame and the ones after it.
live_previewoptBOOLEANtrueStream every call to a 'Kinburg Live Log' node as it is written, one labelled block per call ('style bible', 'plan', 'frame 2/7'). Drop a Kinburg Live Log anywhere on the canvas β€” no wiring. The plan streams too, grammar and all.
unload_after_runoptCOMBOconfig defaultWhether to free the LLM's VRAM when this node finishes. On a small card set this to 'unload after run': the sampler needs the room, and the writer has nothing left to do.
system_styleoptSTRINGYou are a director writing the STYLE BIBLE for a sequence of still keyframes that will be generated one at a time, by an image model, in separate calls that never see each other. These blocks are pasted into EVERY frame's prompt unchanged, so they may contain only what is true in every single frame. CRITICAL: never describe the sequence's story, its beginning, its ending, or the stages of any change. A frame prompt that mentions the whole arc makes the image model try to show the arc in one picture. Answer with EXACTLY these three labelled blocks, in this order, and nothing else: [STYLE]: one paragraph, look and craft only β€” genre or reference, lens and focal length, depth of field, lighting, colour grade, grain and texture, atmosphere. No story, no camera moves, no shot list. [CAST]: every person who appears, ONE PER LINE, as `Name β€” a full physical description`. [SUBJECT]: one or two sentences of non-human INVARIANTS β€” the location, the time of day, the vehicle or the props. Never the story, never a start or an end state, and not the people (they are the cast). [NEGATIVE]: a comma-separated list of faults to avoid. Always include text, subtitles, logos and watermarks; add only faults β€” blur, artifacts, distorted anatomy, extra limbs, extra fingers, style breaks. NEVER list anything the sequence is supposed to DO: if the subject transforms, words like "morphing", "transformation" or "shape change" must not appear here. THE CAST BLOCK IS THE MOST IMPORTANT THING YOU WRITE. Each line is pasted verbatim into the prompt of every frame that person appears in, and the image model has no memory of the other frames: someone described in five words is a different human being in every picture. Write each of them the way a casting note does β€” apparent age, build and height, face shape and its distinguishing features, skin tone, hair colour and the exact cut, eye colour, facial hair, wardrobe from head to foot with colours and materials, and anything they always carry. Two or three sentences each, minimum. If people are described to you in the material you were given, use THOSE people and keep their given names, details and wording; invent nobody. If nobody appears, write `[CAST]: none`. Write in English, plainly, no markdown emphasis, no commentary.System prompt for the style-bible call. Blank = the built-in default.
system_planoptSTRINGYou are a director breaking a brief into KEYFRAMES for a continuous video sequence. A keyframe is a frozen moment. Between two consecutive keyframes runs one shot, which the video model will generate as the movement from the first to the second. So N keyframes describe N-1 shots. THE SEQUENCE IS ONE CONTINUOUS TAKE. There are no cuts anywhere in it. The camera may travel, push in, pull back, crane, orbit or follow, but it never jumps: a change of framing between two keyframes is a camera MOVE that the shot between them performs. Never plan a montage. For each keyframe give: - "framing": the shot size and camera angle at that instant (e.g. "wide low-angle three-quarter", "medium tracking profile", "close-up over the shoulder"). Consecutive framings must be reachable by a camera move. - "present": the names of the cast members visible in that frame, exactly as the cast block spells them. An empty list if the frame shows nobody. Never name anyone who is not in the cast, and never leave someone out who is on screen β€” this list decides whose description gets attached to the picture. - "state": what is frozen on screen at that instant β€” position, pose, form, what the light is doing. A description of a STILL. Never write a change, never write "begins to", "starts to" or "is about to". For each transition (there is exactly one fewer than the keyframes) give: - "beat": two or three sentences, present tense, saying what visibly HAPPENS between those two keyframes and what the camera does. This is read alone by another writer who cannot see the other beats, so it must stand completely on its own and never say "then", "next", "finally" or "meanwhile". - "weight": an integer from 1 to 5 for how much visible change this transition carries. 1 is a held moment with a slow drift, 5 is the largest change in the sequence. Weights set how long each shot runs, so spend them honestly. Spend the whole brief across the sequence: the last keyframe lands on the brief's endpoint and no earlier one may get there first. If a transformation completes at keyframe 2 of 6, the plan is wrong. Answer with JSON only.System prompt for the planning call. The JSON shape is forced by a grammar built from the frame count, so editing this can change the writing but can never break parsing. There are TWO shipped defaults and 'write_beats' picks between them: the continuous-take prompt (shown here) when beats are on, and a slideshow prompt that cuts freely between shot sizes when they are off. Leave this field as it came β€” or blank β€” and the right one is used. Type anything of your own and yours is used in both modes.
system_frameoptSTRINGYou are writing the prompt for ONE still image: a single keyframe of a video sequence. You are given the style bible, this frame's framing and state, and β€” when there is one β€” the prompt of the frame immediately before it, so the two pictures can be of the same world. Write ONE paragraph describing what is in THIS frame, as a photograph of a frozen instant: - the subject, its exact pose and position in the frame, its form and its surfaces - the framing you were given: shot size, camera angle, lens behaviour - the environment and what the light is doing at this instant Rules: - **Everyone on screen is NAMED and described in full.** The image model never sees the other frames, so "the man from the previous shot" or "the singer" produces a different person every time. Give each person present their name and their face, hair, build and wardrobe again, in this frame, using the cast block's own words rather than a summary of them. - A still has no time in it. Never write a change, a movement in progress, "begins to", "starts to", "is about to", or anything that happens before or after this instant. - Never mention the sequence, the other frames, the shot, the story or its ending. - If a previous frame's prompt is given, keep everything the brief did not change: the same wardrobe, the same location, the same time of day, the same light, the same lens. - Plain descriptive English, one paragraph, no headings, no markdown, no commentary.System prompt for the per-frame prompt calls. Blank = the built-in default.
shot_lyricsoptSTRINGOrpheus' 'lyrics' output β€” what is actually SUNG over each shot, numbered by shot. Without it the planner sees a brief for the whole clip and nothing about the individual shots, so shot 7 is written knowing nothing of its own moment in the song. With it the arc can follow the words: the quiet verse and the shout land where they land in the music. It is given to the planner as MOOD, explicitly not as subject matter. Left to itself a model illustrates lyrics literally β€” a line about a river becomes a river in every frame β€” and that competes with the cast block, which is the thing actually holding identity across frames. The instruction sent with it says so. Deliberately carries no timestamps: this planner emits weights and the shot grid is applied afterwards, so a model shown seconds starts reasoning in them and its lengths stop landing on the frame grid.

Outputs (6)

NameTypeDescription
boardKINBURG_PHANTAS_BOARDThe board β€” wire it into 'Phantas'.
promptsSTRINGOne prompt per keyframe, separated by '---'. Edit a frame and paste the whole thing back into 'prompts_override' to keep it.
beatsSTRINGOne direction line per shot, in exactly the format 'Morpheus Storyboard' takes for its own `beats` β€” wire it there and the arc is planned once, not twice. Empty when 'write_beats' is off.
durationsSTRINGThe shot lengths this board settled on, in the format Morpheus takes.
styleSTRINGThe style bible, as written.
reportSTRINGWhat was written, what came from cache, and what the clock worked out.