LTX-Video 2.5 Video Prompt Builder (Official Encyclopedic LTX)
LTX-Video Prompts With Dialogue, SFX, and Multi-Shot Cuts
- video_prompt
- audio_cues
Video models are unforgiving about vague prompts. "A woman in a cafe" gets you a smeary pan with no idea what anyone's doing; LTX especially rewards camera detail and explicit action. This node is the pack's answer to that: it assembles a full descriptive video_prompt and a separate audio_cues string - dialogue tags, SFX brackets, ambient sound - so one node feeds both the visual and the audio side of LTX's multimodal pipeline.
It comes from the ComfyUI-StudioPromptDirector pack, and it's aimed at the LTX-Video 2.x family, which is notable for being the first open model to generate synchronized video and audio in one pass. One honesty note before you get excited: the name says "LTX-Video 2.5," but as of mid-2026 no 2.5 checkpoint has shipped - LTX 2.3 is the current open release, and 2.5 (a promised new latent space) was still roadmap at the January AMA. The node's output syntax works fine with the 2.x family and its IC-LoRAs, so don't let the version sticker put you off; just don't go hunting for a model called 2.5.
What it builds
The main output, video_prompt, is a comma-joined shot description assembled from your choices:
sequence_structure- the headliner. Six modes including Single Continuous Take, Multi-Shot Scene (2–4 cuts in one prompt), Dub-It IC-LoRA (speech/lip-sync replacement on existing footage), Video-to-Video Edit IC-LoRA, time-lapse, and slow-motion.camera_motion_take1- 16 choices from slow push-in to whip pan to SnorriCam rig.scene_pacing,motion_choreography_preset, andmulti_shot_cut_transition- the transition enum is where the cinematic grammar lives: "A hard cut transitions to…", "A match cut connects…", "A sudden smash cut jumps to…".- The labeled boxes:
Primary_Actor_Action_Take1,Secondary_Actor_Reaction_Take1,MultiShot_Cut2_Action_and_Framing, andSynchronized_Dialogue_English.
The second output, audio_cues, is the part most prompt builders don't have. It collects your dialogue as an <d>[English] "…"</d> tag, your SFX as [sharp_slap_sfx], [gasp], [cafe_ambience]-style brackets, plus the ambient_soundscape_preset. That's the LTX audio side's language, and it's also genuinely what its Dub-It and speech-replacement LoRAs want to see.
Inputs you'll actually touch
For a beginner, three things get you 90% of the way: sequence_structure (pick Single Continuous Take for your first tries), camera_motion_take1, and the dialogue box. The defaults are already a working coffee-shop argument scene, which is nice for seeing the output format without building anything. Everything else is flavor you dial in once the format clicks.
Install and troubleshooting
Pure Python, zero dependencies, no models to download - the pack's requirements.txt is literally a comment line.
cd ComfyUI/custom_nodes
git clone https://github.com/nexusfinancial-dev/ComfyUI-StudioPromptDirector.git
Or search "ComfyUI-StudioPromptDirector" in ComfyUI Manager and restart. Nodes land under the StudioPromptDirector category.
Where people get burned: LTX is a great model that's historically been finicky about adherence. Community reports are consistent - if you don't specify camera details and quality keywords, you get visibly lower quality output, so keep the camera and choreography dropdowns populated rather than leaving defaults. Audio artifacts (mumbling, dropouts) were a known LTX-2 era complaint that the 2.3 release and its distilled LoRA fixed - if you're chasing clean dialogue, use the distilled path and treat the first few runs as an experiment, not a verdict. And remember this node only writes text: video_prompt goes into the CLIP Text Encode of your LTX workflow, audio_cues into whatever audio/text condition your setup uses. It won't make LTX obey - but it gives the model a shot it can actually understand.
Inputs (12)
| Name | Type | Default | Description |
|---|---|---|---|
| sequence_structure | COMBO | Single Continuous Take (Unbroken Cinematic Camera & Motion) | 6 options: Single Continuous Take (Unbroken Cinematic Camera & Motion), Multi-Shot Scene (2-4 Cuts Joined by Hard Cuts in One Prompt), Dub-It IC-LoRA (Speech & Lip-Sync Replacement on Existing Video), Video-to-Video Edit IC-LoRA (Additive Element & Style Replacement), Dynamic Time-Lapse / Hyper-Lapse Accelerated Sequence, Extreme Slow-Motion Phantom Take (1000 FPS High-Speed Feel) |
| camera_motion_take1 | COMBO | Slow Push-in Dolly (Dramatic Tension Build) | 16 options: Slow Push-in Dolly (Dramatic Tension Build), Slow Pull-out Dolly (Revealing Scene Context & Scale), Static Locked Tripod (Clean Dialogue & Performance Focus), Dynamic Tracking Shot (Following Actor Motion Smoothly), Smooth 360-Degree Arc Orbit (Heroic & Cinematic), Subtle Handheld Movement (Documentary Realism & Urgency), +10 |
| scene_pacing | COMBO | Tense Heated Verbal Confrontation (Medium-Fast Pacing) | 7 options: Tense Heated Verbal Confrontation (Medium-Fast Pacing), Climactic Physical Action Impact (Fast & Sudden Shock), Slow Emotional Cinematic Pacing (Deliberate & Heavy), Suspenseful Building Tension (Gradual Acceleration), Fast-Paced Action Chase (Urgent & Kinetic), Gentle Serene Contemplation (Fluid & Calm), +1 |
| motion_choreography_preset | COMBO | The Climax Slap & Recoil Reaction Shock (Hand on Flushed Cheek) | 12 options: The Climax Slap & Recoil Reaction Shock (Hand on Flushed Cheek), Violent Table Slam with Trembling Cups & Coffee Spilling, Heated Argument with Furious Accusing Finger Pointing, Standing Up Rapidly from Chair in Outraged Confrontation, Devastated Tears Rolling Down Cheek in Emotional Close-Up, Subtle Micro-Expressions, Blinking & Articulate Lip Sync, +6 |
| multi_shot_cut_transition | COMBO | A hard cut transitions to a medium close-up of | 7 options: A hard cut transitions to a medium close-up of, A match cut connects the physical action to, The camera whip-pans across to reveal, The frame dissolves softly into a wide shot of, A sudden smash cut jumps to an extreme close-up of, An over-the-shoulder cut reveals the shocked face of, +1 |
| ambient_soundscape_preset | COMBO | Modern Luxury Cafe (Coffee Grinder, Cup Clinking, Soft Murmur) | 8 options: Modern Luxury Cafe (Coffee Grinder, Cup Clinking, Soft Murmur), Quiet High-Tension Room (Audible Breathing & Ticking Clock), Heavy Rainstorm & Distant Thunder Outside Windows, Crowded Restaurant (Lively Background Chatter & Cutlery), Sudden Dramatic Silence with Deep Sub-Bass Drone, Bustling Neon City Street (Cars Hissing on Wet Asphalt), +2 |
| 🏷️_Primary_Actor_Action_Take1opt | STRING | Maya stands up aggressively with her right arm swinging forward | — |
| 🏷️_Secondary_Actor_Reaction_Take1opt | STRING | Mariam recoils in shock, clutching her cheek as tears well up in her eyes | — |
| 🏷️_MultiShot_Cut2_Action_and_Framingopt | STRING | close-up of Mariam's flushed red cheek, her hand trembling. The cafe murmur continues across the cut. | — |
| 🏷️_Synchronized_Dialogue_Englishopt | STRING | Maya whispers furiously, "I told you never to say that again!" | — |
| 🏷️_Sound_Effects_Tagsopt | STRING | [sharp_slap_sfx], [gasp], [cafe_ambience] | — |
| 🏷️_DubIt_Native_Script_Replacementopt | STRING | — |
Outputs (2)
| Name | Type | Description |
|---|---|---|
| video_prompt | STRING | — |
| audio_cues | STRING | — |