⭐ Star LTXV 2.5 All-in-One (BETA)
LTX 2.5's whole pipeline in one node — video, audio, first/last-frame, and a BETA label you should believe
- first_frame
- last_frame
- audio
- model_override
- sound_settings
- preview
- images
- audio
- frame_rate
- latent
- model
- clip
- vae
- audio_vae
The name says "All-in-One" and it means it. Star LTXV 2.5 All-in-One (BETA) is the StarNodes port of the official LTX 2.5 workflow templates, compressed into one node that handles text-to-video, image-to-video, image-plus-audio, first-and-last-frame keyframing, and even an audio-only mode - plus the two-pass render pipeline, model/LoRA/CLIP/VAE caching, and internal sound processing. If you want to run LTX 2.5, this is the fastest way from zero to a finished clip that the pack offers.
It's in ⭐StarNodes/Video, and it's the pack's newest and most ambitious node - which is also why the BETA label is honest, not decorative.
How it works - the modes
The mode dropdown is where you start, and it decides the whole graph:
- text_to_video - prompt only.
- image_to_video - connect
first_frame. - image_audio_to_video - connect
first_frameand anaudiofile; the audio is trimmed to the video length and preserved. - first_last_frame_to_video - connect
first_frameandlast_frame; runs a single full-resolution pass with both frames added as keyframe guides (LTXVAddGuide, same approach as the official FLF2V template). This is the one mode that skips the two-pass structure. - audio_only - no real video at all; a single 30-step pass at 64×64 whose only output that matters is the audio. Yes, that's as odd as it sounds, but it's the official workflow's approach and it works.
The main modes (T2V / I2V / I2V+audio) use the two-pass pattern: pass 1 at half resolution, a 2× latent upscale in between, pass 2 at full resolution for the refine. LTX 2.5 uses a single text encoder (a ltxv-type Gemma encoder), which is a real simplification versus the older two-encoder setups.
The inputs that matter
base_model,clip_1,vae,audio_vae,upscale_model- the internal loaders. Defaults point at LTX 2.5 files (ltx-2.5-22b-distilled-transformer-comfy-int8-convrot, thegemma4-12bencoder,ltx-2.5-video-vae-conv-bf16, etc.) from the standard model folders. Each is cached and only re-loaded when you change the selection.sigma_preset- the baked 8/12/16-step schedules from the original workflow, plus plain 20/30/40/50-step sampler options, orcustomwithcustom_sigmas_pass1.cfg(default 1.0) and the two sampler dropdowns - both passes.sound_settings- optional; plug in the pack's Star Video Sound Enricher Option bundle and the audio output gets cleaned and enriched (at least 44.1 kHz, never downsampled) before it leaves the node.override_audio- the subtle one. By default (in T2V/I2V/FLF modes) a connectedaudioinput is ignored and the model-generated audio goes to the output; flip it on to pass your own audio through instead. Inimage_audio_to_videomode the connected audio always passes through.
Outputs
images, audio, and frame_rate - the same trio as the LTX 2.3 node, and the audio output is the part most people underestimate. LTX 2.5 generates synchronized audio, and this node decodes it from the first (high-step) pass, which is the version with the best audio quality.
Installing it
ComfyUI Manager → search Starnodes → Install → restart, or:
cd ComfyUI/custom_nodes
git clone https://github.com/Starnodes2024/ComfyUI_StarNodes
cd ComfyUI_StarNodes
pip install -r requirements.txt
The honest BETA warning
Take the label seriously. LTX 2.5 is brand-new territory (the pack's own README calls this node the newest thing in the pack), and this node wraps a huge amount of machinery - five internal model loaders, two-pass sampling, keyframe guides, an audio path, vendored fallback code for older ComfyUI. When it works it's remarkable; when it doesn't, the single-node black box makes debugging harder than a normal graph would be. Your first three problems will be: (1) missing model files, because the dropdowns only populate from the exact models/ subfolders; (2) an outdated ComfyUI core - this node leans on recent comfy_extras.nodes_lt internals, so update ComfyUI before blaming the pack; (3) VRAM, because a 22B transformer plus a 12B text encoder is a serious ask regardless of how convenient the node is. Update everything, check the console, and budget your card's memory - then enjoy what is genuinely the easiest way to run 2.5.
Inputs (35)
| Name | Type | Default | Description |
|---|---|---|---|
| mode | COMBO | ▶️ image_to_video | text_to_video: prompt only. image_to_video: connect first_frame. image_audio_to_video: connect first_frame AND an audio file. first_last_frame_to_video: connect first_frame AND last_frame (single full-res pass with keyframe guides). audio_only: no real video - one 30-step pass at 64x64, only the audio output matters. |
| positive_prompt | STRING | What you want to see. LTXV likes detailed, film-style descriptions with timestamps. | |
| negative_prompt | STRING | console game, video game, cartoon, childish, ugly | What to avoid. Default is the negative prompt from the original workflow. |
| base_model | COMBO | ltx-2.5-22b-distilled-transformer-comfy-int8-convrot.safetensors | LTXV 2.5 A/V checkpoint from models/diffusion_models. Reloaded only when the selection changes. |
| clip_1 | COMBO | gemma4-12b-with-proj-ltx-2.5-comfy-int8-convrot.safetensors | Single LTXV 2.5 text encoder from models/text_encoders (type ltxv). |
| vae | COMBO | ltx-2.5-video-vae-conv-bf16.safetensors | Video VAE from models/vae (e.g. ltx-2.5-video-vae-conv-bf16). |
| audio_vae | COMBO | ltx-2.5-audio-vae-bf16.safetensors | Audio VAE from models/vae (e.g. ltx-2.5-audio-vae-bf16). |
| upscale_model | COMBO | ltx-2.5-latent-spatial-upscaler-x2-bf16-1.0.safetensors | Latent upscaler from models/latent_upscale_models (e.g. ltx-2.5-latent-spatial-upscaler-x2). Used between the passes. |
| video_size | COMBO | HD | HD ~1280px, FHD ~1920px (same tables as the Star LTX Video Settings node), Custom = custom_width/height below. |
| ratio | COMBO | 1:1 | Aspect ratio. Overridden by the input image's ratio when 'ratio_from_image' is enabled and an image is connected. |
| ratio_from_image | BOOLEAN | true | Pick the closest preset ratio to the connected image. Falls back to 'ratio' when no image is connected. |
| custom_width | INT | 102432–8192 | Only used when video_size = Custom. |
| custom_height | INT | 102432–8192 | Only used when video_size = Custom. |
| frame_rate | INT | 241–120 | Frames per second of the output video. |
| seconds | INT | 101–120 | Video length in seconds. Frame count is snapped to 8n+1 (4s @ 25fps = 97 frames). |
| seed | INT | 00–18446744073709550000 | Shared by both sampling passes. |
| sigma_preset | COMBO | Main-pass noise schedule. 8/12/16 = the baked schedules from the original workflow's note node (12 = default, 8 = faster, 16 = finer). 20/30/40/50 = plain sampler steps with the normal scheduler, no custom sigmas. 'custom' uses custom_sigmas_pass1 below. | |
| first_frameopt | IMAGE | Start frame / guide image (image_to_video, image_audio_to_video and first_last_frame_to_video modes). | |
| last_frameopt | IMAGE | Last frame (first_last_frame_to_video mode only). Center-crop resized to the video size and added as the final keyframe. | |
| audioopt | AUDIO | Voice / music track (image_audio_to_video mode only). Trimmed to the video length and preserved as-is. Ignored in all other modes. | |
| lora_1opt | COMBO | Optional LoRA stack, applied in order 1 -> 3. | |
| lora_1_strengthopt | FLOAT | 0.60-100–100 | The distilled LoRA in the original workflow ran at 0.6. |
| lora_2opt | COMBO | 1 options: None | |
| lora_2_strengthopt | FLOAT | 1.00-100–100 | — |
| lora_3opt | COMBO | 1 options: None | |
| lora_3_strengthopt | FLOAT | 1.00-100–100 | — |
| custom_sigmas_pass1opt | STRING | 1.0, 0.995833, 0.991667, 0.9875, 0.983333, 0.979167, 0.975, 0.93125, 0.847917, 0.725, 0.522917, 0.28125, 0.0 | Only used when sigma_preset = custom. |
| sigmas_pass2opt | STRING | 0.85, 0.7250, 0.4219, 0.0 | Second-pass (refine) schedule. Default from the workflow. |
| cfgopt | FLOAT | 1.00–100 | Both passes. 1.0 for distilled models, as in the workflow. |
| sampler_pass1opt | COMBO | euler_ancestral | Sampler for pass 1 (half resolution). |
| sampler_pass2opt | COMBO | euler_ancestral | Sampler for pass 2 (full resolution refine). |
| weight_dtypeopt | COMBO | Override base-model dtype. 'default' = as stored. | |
| model_overrideopt | MODEL | Optional external model (e.g. patched with flash/sage attention). When connected, this is used instead of loading 'base_model' from the dropdown, and the LoRA stack below is applied to it directly. | |
| sound_settingsopt | SOUND_SETTINGS | Optional sound processing from a 'Star Video Sound Enricher Option' node - the audio output is cleaned up and enriched with these settings (at least 44.1 kHz, never downsampled) before it leaves the node. | |
| previewopt | STAR_PREVIEW | Optional live sampling preview from a '⭐ Star Preview' node - while this node is sampling, an animated preview of the video latent is shown on the Star Preview node (fixed: 512 px, quality 80, 8 fps). |
Outputs (8)
| Name | Type | Description |
|---|---|---|
| images | IMAGE | — |
| audio | AUDIO | — |
| frame_rate | FLOAT | — |
| latent | LATENT | — |
| model | MODEL | — |
| clip | CLIP | — |
| vae | VAE | — |
| audio_vae | VAE | — |