Nodes/ComfyUI-PainterNodes/Painter LTX2V Omni
ComfyUI Node

Painter LTX2V Omni

The node that gives Bernini a voice

By princepainter·Created 8 months ago·Updated 4 days ago· 195
Painter LTX2V Omni
  • positive
  • negative
  • video_vae
  • audio_vae
  • start_image
  • end_image
  • source_video
  • reference_image1
  • reference_image2
  • positive
  • negative
  • video_latent
  • audio_latent
◄width768►
◄height512►
◄length97►
◄frame_rate25.00►
◄strength0.70►
◄reference1_frame_idx0►
◄reference2_frame_idx0►

Bernini makes gorgeous silent video, and "the king is mute" is the community's favorite jab at it. The standard workaround is to run the render through LTX-2.3 to add sound - and PainterLTX2Vomni is the node that makes that one pass instead of a Franken-workflow. It's an omni (video + audio) conditioning and latent builder for LTX 2.x: feed it the silent video, prompt conditioning, both VAEs, and it returns the video latent, the audio latent, and conditioning for a single LTX omni sampling pass that outputs picture and synced sound.

What it is

A model/conditioning/ltxv node, and the newest star in the pack's changelog: "can dub and lip-sync Bernini-generated silent video, works very well, worth trying." The author frames it as supporting text-to-video, image-to-video, first/last-frame, and reference video tasks too - but the reason people install it is the Bernini dubbing path.

How it works

LTX-2/2.3 is a multimodal model: one DiT renders both video and audio, and it has two VAEs - a video_vae for the frames and an audio_vae for the sound. This node expects both (they come from your LTX 2.3 loader), then builds both latents. The outputs tell the story: video_latent and audio_latent - you feed both into the omni sampler, and the decoded result is video with a native audio track, mouth movements included. That's the whole magic of the "dub a silent Bernini clip" workflow: the LTX pass re-renders with the audio baked in, giving you Wan-2.2-level motion (from Bernini) plus LTX's sound.

The strength input (default 0.7) controls how hard the omni pass re-conditions the source - lower keeps the original footage closer, higher lets LTX re-animate more aggressively for the lipsync. reference1_frame_idx / reference2_frame_idx pick which frames of the source get treated as reference anchors (negative values work as offsets). Resolution steps by 32 and length steps by 8, matching LTX-2.3's hard rule that width/height are divisible by 32 and frame count is 8n+1 - off-grid values fail on this model, so the stepping is doing you a favor.

Inputs and outputs that matter

  • source_video - the silent clip to dub (this is the Bernini output slot).
  • video_vae / audio_vae - both VAEs from your LTX 2.3 stack.
  • strength - how much the omni pass re-animates the source.
  • frame_rate / length - timing and frame count (8n+1).
  • start_image / end_image / reference_image1/2 - optional anchors for I2V and reference tasks.
  • Outputs: positive / negative / video_latent / audio_latent → into the omni sampler.

Install

Part of the PainterNodes pack. ComfyUI Manager → search "PainterNodes" → install → restart, or:

cd ComfyUI/custom_nodes
git clone https://github.com/princepainter/ComfyUI-PainterNodes
cd ComfyUI-PainterNodes
pip install -r requirements.txt

The full graph is in workflows/bernini+ltx2.3.json - that's the reference for the "Bernini render → LTX dub" pipeline.

Common issues

  • No audio in the output - you decoded only the video latent. The omni sampler needs both latents; a sampler that doesn't know about the audio latent will cheerfully drop it.
  • Resolution/length errors - LTX 2.3 enforces the 32 / 8n+1 rules hard, and off-grid values fail rather than round. The node steps by 8 on length, but if you're typing odd numbers yourself, expect errors.
  • It's heavy - LTX 2.3 is 22B. On 16GB cards the GGUF quants are the way; on 8–12GB you're in slow territory and 64GB system RAM helps more than you'd think.

This is the node that patches the single most complained-about gap in the best open editing model of 2026. If you run Bernini, you'll want it.

Categorymodel/conditioning/ltxv

Inputs (16)

NameTypeDefaultDescription
positiveCONDITIONING—
negativeCONDITIONING—
video_vaeVAE—
audio_vaeVAE—
widthINT76864–16384—
heightINT51264–16384—
lengthINT971–16384—
frame_rateFLOAT25.000–1000—
strengthFLOAT0.700–10—
reference1_frame_idxINT0-9999–9999—
reference2_frame_idxINT0-9999–9999—
start_imageoptIMAGE—
end_imageoptIMAGE—
source_videooptIMAGE—
reference_image1optIMAGE—
reference_image2optIMAGE—

Outputs (4)

NameTypeDescription
positiveCONDITIONING—
negativeCONDITIONING—
video_latentLATENT—
audio_latentLATENT—