PromptStudio_Video
Structured video direction and MiniMax H3 generation for ComfyUI Prompt Studio
Nodes (5)
The context-splicing node
The bookmark node that lets MiniMax H3 renders continue later
Cutting the repeated head out of a continued MiniMax H3 clip
This node builds the conditioning from a shot list
4-step MiniMax H3 without memorizing the sigma-shift table
PromptStudio_Video
Build MiniMax H3 videos in a visual, shot-based studio that keeps timing, references, dialogue, camera, and sound under your control.
PromptStudio_Video is the standalone video companion for ComfyUI_PromptStudio and works directly with current native ComfyUI MiniMax nodes.
Direct the video on a real timeline
Plan multiple shots on a proportional, frame-snapped timeline. Reorder and trim them visually, then edit each shot's setup, camera, sound, visible text, and chronological action or dialogue steps.
Each Shot Editor also contains a nested, shot-local timeline for overlapping actions, dialogue, generated sound cues, and exact imported audio. Semantic items use the timeline to compile guide-native chronological phrases such as "immediately," "then," and "while"; MiniMax receives numeric timestamps only for shot cuts. Exact audio is mixed after generation, so its waveform, gain, fades, source trim, and placement remain deterministic without forcing the project into REF2VA mode. The backend probes the actual audio stream duration through PyAV and supports its common audio containers and codecs, including WAV, MP3, FLAC, OGG/Opus, M4A/AAC, AIFF, and WMA where the installed FFmpeg build provides the decoder.
<!-- Hero screenshot slot: full Video Studio with preview, timeline, and shot navigator. Suggested file: docs/images/video-studio-timeline.png  -->Collaborate with two local AI Directors
Ask the Video director to critique or compose the whole production, or use the Shot director for a focused change. Both return reviewable structured proposals; nothing changes until you explicitly apply it.
<!-- Screenshot slot: Video director proposal with apply/discard review. Suggested file: docs/images/grand-director.png  -->Use every MiniMax H3 generation mode from one project
Drop in first and last frames, images, video, or audio, assign clear semantic roles, and let the project route T2VA, I2VA, FL2VA, L2VA, or REF2VA automatically. Reference labels, dialogue, lyrics, and visible text remain explicit and protected.
<!-- Screenshot slot: media library with several references and MiniMax roles. Suggested file: docs/images/reference-modes.png  -->Preview exact prompts and replay exact generations
The structured project document remains authoritative while a deterministic compiler builds the guide-compliant MiniMax prompt. Every result keeps its prompt, project state, routing, and executable workflow snapshot for inspection and replay.
Completed renders also expose Continue video. Video Studio sends the final 39 frames and their exact 65 audio-latent steps into the target head of a new MiniMax H3 sample. ComfyUI's native nested denoise mask protects the complete picture prefix and the first 57 audio steps; the final eight audio steps use a half-cosine release into the generated future. Native guides inside the protected prefix are removed so they cannot compete with the exact target latent.
After decoding, Video Studio saves both a separately playable trimmed extension and a private overlap-bearing assembly segment. Cumulative output uses the full 39-frame visual blend and gives the incoming segment ownership of overlap audio, so the model's Soft AV audio release reaches the final soundtrack. Every generation stores a compact v3 latent tail plus immutable public and assembly lineage. Parent generations are never overwritten, so any shorter version can still be played or branched. Generated parents reuse their sampled latent directly; older or imported parents fall back through the audiovisual VAEs. No third-party motion-context node pack or ComfyUI runtime patch is required.
The continuation dialog offers two paths. Quick generate sends one concise free-form development directly into a single continuation shot. Build full extension creates a durable extension project linked to the immutable source render and workflow snapshot. The extension uses the normal timeline and Shot Editor, so it may contain multiple cuts, ordered action and verbatim dialogue steps, camera changes, visible text, generated sound cues, ambience, and music. Full planning runs as a background job with progress, cancellation, retry and an explicit Apply extension plan step. Applying a plan creates one child project and preserves the parent; a failed or cancelled plan is not silently applied. The same provider scheduler used by Image and the Video Directors owns planning. Its authored timeline describes only the new 5–15 second tail; Video Studio injects the 39-frame Soft AV handoff, offsets authored cuts and first-shot event timing, then removes the overlap automatically. Shot director works on every extension shot and receives a read-only description of the source's final shot so Shot 1 continues rather than replaying the boundary. Video director and new media references are intentionally disabled in structured extensions for now.
Every completed, failed, interrupted, or cancelled continuation also exposes Regenerate extension. It queues only a new tail from the same immutable parent handoff; the original video and earlier accepted segments are not sampled again. Quick extensions may revise their free-form development and duration in the dialog, while structured extensions use the project's current authored shots. A fresh seed is used, the replacement is assembled against the same saved source segments, and the previous variation remains available for comparison.
<!-- Screenshot slot: compiled prompt preview beside generation history and replay. Suggested file: docs/images/prompt-preview-and-replay.png  -->Move seamlessly from still image to video
With the companion image extension installed, send a Prompt Studio result straight into the active video project and assign it as a frame or reference without exporting and re-importing it manually.
<!-- Screenshot slot: imported Prompt Studio image in the Video Studio media library. Suggested file: docs/images/image-studio-handoff.png  -->This repository is under active development. The sections below cover requirements, the Director document, AI behavior, workflows, integration, and development details.
Requirements
- The companion
ComfyUI_PromptStudiorepository, which owns shared local-LLM settings and Llama.cpp server management, including optional startup with ComfyUI and its delayed shutdown watchdog when ComfyUI exits. - A current ComfyUI build containing the native MiniMax H3 and video nodes.
- MiniMax H3 model, text encoder, and VAE files.
- KJNodes is recommended for the optimized workflow path. It is not imported by the Director itself.
No additional pip package is currently required. ComfyUI supplies PyTorch, PyAV, torchaudio, Pillow, and its native media loaders.
The installation metadata retains Python >=3.9. Only Python 3.14 through
the ComfyUI VENV was verified locally for this audit; Python 3.9–3.13 remain
unverified. The companion's quality workflow targets 3.14, and its
release checklist
separates unit/browser-fixture evidence from native GPU and media validation.
There is no unconditional startup installer for host-provided packages.
Runtime compatibility is capability-based: the host must provide native
MiniMaxH3ImageToVideo, MiniMaxH3ReferenceToVideo, SaveVideo, H3 nested latent
and denoise-mask support, the required media codecs, and isolated graphToPrompt
conversion with serialized subgraphs. /promptstudio-video/capabilities
advertises native_h3_soft_av_39; this is the supported 39-frame continuation
contract, not proof that a model or GPU render has been exercised on that host.
The two default workflows are maintained as readable source JSON in
workflows/. video/default_setup.py loads those files directly
with the standard library; there are no opaque compressed workflow strings to
regenerate. Source serialization is deterministic, and packaging tests check the
normal/Turbo node contracts and round-trip data. Keep this directory in source
and registry distribution archives.
Installation
Install the repository as:
ComfyUI/custom_nodes/PromptStudio_Video
Restart ComfyUI after installation.
Current capabilities
The current vertical slice contains:
Prompt Studio MiniMax H3 Director, a ComfyUI custom node;- automatic T2VA, I2VA, FL2VA, L2VA, and REF2VA routing;
- a deterministic prompt compiler following MiniMax's official base and full-reference formats;
- native ComfyUI MiniMax conditioning and latent creation;
- lazy FL2VA/REF2VA model selection;
- first-frame, last-frame, image, video, and audio reference loading;
- nested per-shot action, dialogue, generated-sound, and exact-audio timelines with overlapping ranges;
- post-render exact-audio overlay/replacement with source trimming, gain, fades, and authoritative duration probing;
- dependency-free native H3 Soft AV continuation with exact target-prefix masks, immutable parent/child lineage, separately viewable extension segments, and overlap-aware cumulative outputs;
- a transactional full-screen Director Canvas with frame-snapped shot trimming, drag-to-reorder editing, media drag-and-drop, camera direction, dialogue, visible text, sound, and media roles;
- a standalone Video Studio with durable projects, authoritative manual shot editing, a docked proportional timeline, frame-snapped trimming, per-shot editors, draggable action/dialogue steps, media roles,
[PSV]workflow discovery, deterministic prompt preview, queueing, progress, video history, and exact workflow replay; - context-budgeted Video director and Shot director workflows for KoboldCpp, Ollama, or Llama.cpp, with conversational advice, validated structured proposals, stale-document protection, and explicit apply/discard review; and
- a versioned capability and document API for Prompt Studio integration.
Director document
The node stores a versioned JSON document rather than treating the compiled prompt as its source of truth. Its references determine Auto mode:
| Reference roles | Resolved mode |
| --- | --- |
| No references | T2VA |
| One first_frame image | I2VA |
| One first_frame and one last_frame image | FL2VA |
| One last_frame image | L2VA |
| Subject/style/scene/video/audio references | REF2VA |
The final prompt is compiled at execution time. Alignment instructions use the
effective MiniMax duration after the frame count is snapped upward to the
model's 17k+5 temporal grid.
Video Studio records the native dimensions of imported images and videos and
offers each aspect ratio in the Canvas selector. First-frame, last-frame, video
editing, and video continuation roles can link the canvas to their source
geometry. The linked ratio is resized to a configurable target megapixel value
(about 1.03 MP by default, matching 768 × 1344) and both axes are rounded to
MiniMax H3's required multiple of 32.
Each shot owns an authoritative chronological steps list. Action and dialogue
steps can be interleaved and are compiled in their exact top-to-bottom order;
dialogue text remains verbatim inside MiniMax <d> tags. Normalized documents
contain no parallel action or dialogue mirrors. Older saved documents receive
a one-way conversion to steps when loaded. Camera movement is stored as
motion type, amplitude, speed, and target, then compiled into natural English.
Director proposals write performance changes only through steps; legacy shot
sequence fields are rejected. Synchronized shot sounds remain separate
from the chronological performance sequence, while the overall soundscape and
non-diegetic music keep their guide-defined roles. Complete silence
suppresses dialogue, shot sounds, ambience, and music in the compiled prompt
without deleting the editable document content.
The editable Production brief is a planning synopsis for the user and Video director. It is not compiled into the MiniMax prompt. The Video director uses it to infer the useful shot count and translates its visual beats into the timeline-specific shot fields. It is also available as Production brief (planning only) in the workflow node's Director Canvas.
Reference and shot identifiers use MiniMax's guide-exact grammar:
<Picture 1>, <Video 1>, <Audio 1>, <Subject 1>, and [Shot 1]. Legacy
compact forms are normalized outside dialogue and visible text, which remain
verbatim.
AI Directors
Use Ask Video director in the production card for whole-video critique and on-demand multi-shot composition. The Video director receives every shot plus the production brief, global audiovisual fields, references, and MiniMax constraints. Approved proposals may update project-level style and sound, reference definitions and retention analysis, revise any shot, change cut times, or add and remove shots. Small changes are field patches that preserve everything not mentioned. Explicit full-production or full-scene rewrites use replacement operations: every omitted old prompt field is cleared, protected text may be replaced as part of that reviewed rewrite, and any stale raw-prompt override is discarded when the proposal is applied. REF2VA proposals must represent every source reference, use every defined label in the summary and applicable shots, and populate all six full-reference sections before generation.
Use Shot director in the shot popup editor for the tighter shot-focused workflow. It receives only the selected shot and immediate neighbors, and its proposals can update descriptive fields, sound, camera, and append newly requested dialogue on that selected shot.
Both buttons open the local LLM consultation panel with separate conversation histories. Prompt Studio is the source of truth for local-provider endpoints, selected Ollama/Llama.cpp models, Llama.cpp reasoning caps, fixed-folder config-profile selection, and managed-server configuration. Video Studio reuses those saved settings and the primary Prompt Studio status/model-management APIs while keeping only its Director message/context assembly and a thin adapter to those shared services here.
Before the main answer, a small zero-temperature structured LLM router classifies
the newest turn as a concrete proposal request, discussion, or clarification.
It reads bounded recent conversation, so declarative motion descriptions,
screenplay-like fragments, polite commands, corrections, and follow-ups are not
limited to a hard-coded verb list. Advice-only turns keep the normal
conversational sampling temperature. Edit requests that require a structured
proposal cap temperature at 0.2 so camera types and the document's other enums
remain stable; automatic contract correction uses 0.0.
Exploratory questions and requests for explanation remain conversational and do not create an Apply action. Direct edit commands, including politely worded requests, produce a validated proposal for explicit review before anything is changed.
The latest Director answer has the same response navigation as image Prompt Studio: move left and right through saved alternatives, or use the right arrow on the newest answer to regenerate the same user turn. Each alternative keeps its own validated proposal and apply/discard state, and failed answers remain regenerable.
The default 8,000-character context budget contains compact production data, canonical reference labels, and at most ten recent conversation messages. The Video director never silently drops shots: increase the context budget if a large production does not fit. Older chat turns are discarded first when that budget is tight; the current instruction is never tail-truncated, because doing so could discard a leading constraint such as "do not." If the complete current instruction cannot fit, the request stops with a clear error. KoboldCpp's reported token count and context length further cap the response when available. Chat history is retained locally but approved document state is the authoritative long-term memory.
The response-token setting defaults to 0 (automatic). KoboldCpp uses the
remaining context up to its Director safety cap, Ollama uses fill-context
generation, and Llama.cpp uses the available server context while keeping
private reasoning separate from the final answer. A positive value remains
available when a hard final-answer allowance is desired. Llama.cpp also reuses
the reasoning cap configured in Prompt Studio's active LLM profile.
Director requests run as background jobs so the browser does not mistake a
long processing operation for a failed HTTP request. Through Prompt Studio's shared
provider service, Video Studio reports live token progress when the selected
provider exposes it: KoboldCpp output is counted through its token endpoint,
while Llama.cpp reads active decoded-token counts from /slots and
distinguishes thinking from final response output when possible.
The request timeout defaults to 600 seconds and can be adjusted in Director
settings.
The compact toolbar status icon covers the selected local LLM and ComfyUI. Green means all services are ready, orange means work is being processed, and red means a service needs attention. KoboldCpp processing can be aborted via its API; Llama.cpp cancellation closes Video Studio's active streams without stopping the server. Start/stop/restart, file browsing, config building, and model selection remain in the primary Prompt Studio UI.
Existing dialogue, visible text, and references remain protected in both scopes. The Director may append newly requested dialogue but cannot rewrite or remove an existing line or speaker ID. The Video director cannot remove a shot containing protected dialogue or visible text. Every proposal is previewed against a document hash, normalized, compiled, and applied only after confirmation.
Concrete requests to create, compose, generate, revise, or apply video content require a structured proposal. If a local model returns only prose or an invalid change set, Video Studio makes up to three automatic correction attempts with a larger response allowance; if those omit a valid change set, the response reports the omission instead of implying it can be applied. If intent routing itself fails, automatic fallback mode preserves and validates any proposal the main Director returns rather than silently deleting it.
Up to four images may be dragged into a Director turn or selected with Add image. Each attachment has an explicit usage:
- Describe only sends the image to the vision model for visible-detail extraction without adding it to MiniMax conditioning.
- First frame, last frame, subject, scene, style, pose, camera, and storyboard
usages add the uploaded image to project references with that semantic role
and expose its canonical
<Picture N>token to the Director.
Each image is sent to its own short, deterministic grounding pass for the
current turn, without the requested action or the other images. The pass
extracts subject candidates, private matching aliases, and structured grounded
appearance attributes, which are cached on the project reference. Subjects in
first/last-frame anchors receive stable <Subject N> tokens too. Before the
larger Director call, natural identifiers are resolved to those tokens and the
private aliases and raw subject observations are removed from its context.
Subject definitions keep a concise source binding and explicitly exclude the
source picture's background, scene, lighting, composition, and camera framing.
Grounded attributes remain private unless the user explicitly requests an
appearance detail or change; they are never pasted automatically into the
compiled prompt. Older turns retain compact metadata and the Director's
textual answer, avoiding repeated vision tokens. KoboldCpp requires an active
vision model with MMProj and Jinja; Ollama must report the vision capability;
Llama.cpp must advertise modalities.vision from /props for the selected
model and be started with the matching MMProj file.
The compiler follows MiniMax's official guides:
Minimal workflow
- Load the MiniMax H3 FL2VA and REF2VA diffusion models.
- Apply any desired KJ optimization to each model before connecting it.
- Load the MiniMax CLIP, video VAE, and audio VAE.
- Connect them to
Prompt Studio MiniMax H3 Director. - Connect the selected
model,positive, andlatentoutputs to the normal guider and sampler path. - Decode video and audio with their matching VAEs.
- Use native
CreateVideoandSaveVideofor output. - Save the workflow with a
[PSV]filename prefix for future Video Studio discovery.
The Director requests only the active diffusion-model input through ComfyUI's lazy-input mechanism.
Separate MiniMax H3 Turbo workflow
Keep the normal full-step workflow unchanged and save Turbo wiring as a second
[PSV] workflow. Connect the Director's selected model, mode, width, and
height outputs to Prompt Studio MiniMax H3 Turbo Profile. Connect the
profile's outputs as follows:
modelto any ordinary contentLoraLoaderModelOnlychain, then through model patches andMiniMaxH3SigmaShift; connect that shifted model to both the guider andBasicScheduler;stepstoBasicScheduler.steps;shift_videoandshift_audiotoMiniMaxH3SigmaShift.
The auto_quality preset selects the FL2VA 768p four-step profile only for its
exact 1344x768 training canvas, the mixed-resolution FL2VA eight-step profile
for other base-mode canvases, and the task-specific REF2VA four-step profile
for reference generation. fast_4step selects the mixed-resolution FL2VA
four-step profile away from 1344x768. T2VA, I2VA, FL2VA, L2VA, and continuation
all use the FL2VA family; REF2VA uses its own LoRA.
Install the Kijai rank-reduced Turbo files under ComfyUI/models/loras before
queueing. The node accepts an exact relative path and can also resolve a moved
file by basename when there is only one match. Its default strength is 0.75.
Do not enable EasyCache or SpectrumApplyMiniMaxH3 in the initial Turbo
workflow; their behavior on a four- or eight-step distilled schedule has not
been validated. Add content LoRAs after the Turbo Profile so they remain
independent of inference-mode routing and sampling policy.
Capability endpoint
Prompt Studio can detect the companion at:
GET /promptstudio-video/capabilities
The current response advertises API version 1, prompt-document version 1, the
[PSV] prefix, the MiniMax H3 adapter, and the standalone Video Studio page.
Open the standalone page at:
http://127.0.0.1:8188/extensions/PromptStudio_Video/prompt_studio_video.html
Video Studio discovers saved ComfyUI workflows whose filenames begin with
[PSV]. A compatible workflow currently needs exactly one executable
PSV_MiniMaxH3Director and one native SaveVideo output.
When Video Studio opens without any compatible workflow, it offers to recreate
the bundled [PSV] MiniMax H3 Normal and Turbo workflows. Before changing
anything, a second confirmation lists every required asset and the remaining
download size. Accepted setup reuses matching files from any configured ComfyUI
model path and downloads the missing FL2VA/REF2VA diffusion models, Qwen text
encoder, video/audio VAEs, and all four Turbo LoRAs. Downloads use resumable
.part files; the completed model replaces its target atomically. The saved
workflows are retargeted when an existing matching asset lives in a different
configured subfolder.
Projects and cached workflow snapshots are stored transactionally in ignored runtime files inside this repository. Every queued generation stores its normalized document, compiled prompt, effective duration, output routing, and complete executable workflow snapshot for exact replay.
Development checks
C:\EasyDiffusion\ComfyUI\venv\Scripts\python.exe -m unittest discover -s tests -v
node --check web\js\promptstudio_video_studio.js
node --check web\js\promptstudio_video_redirect.js
git diff --check