Muse Collective LTX Timeline V3
Keeping the Same Face Across a Whole Timeline
- model
- clip
- audio_vae
- vae
- spatial_upscaler
- bg_audio
- base_model
- face_reference_image
- last_chunk_frames
- audio
- stage1_frames
- seed_hunt_preview_1
- seed_hunt_preview_2
- seed_hunt_preview_3
- seed_hunt_preview_4
- seed_hunt_audio_1
- seed_hunt_audio_2
- seed_hunt_audio_3
- seed_hunt_audio_4
The most frustrating part of long-form video generation is that your character is only consistent until they're not. You're five chunks in, the plot's moving, and suddenly the protagonist has a different nose. MuseDirectorSamplerV3 takes V2's director and adds Face ID - a reference-image identity lock that pulls the sampled face toward a real face you provide, chunk after chunk.
V3 is a WIP fork, and the pack is honest about it: V2's full feature set is intact (Seed Hunt included), and the delta is the face machinery. If identity drift is your number-one enemy, this is the version to try.
How Face ID works
The node uses an identity-overlap conditioning pass (under the hood it drives LTX's LTXIdentityOverlapConditioning, the same identity conditioning Lightricks ships for 2.3) to bias the sampled face toward a reference image. The reference goes in through the optional face_reference_image input - and the tooltip gives you the good advice: feed it a close-up face crop, not a full-body shot. The MuseFaceLock node in this same pack exists precisely to produce that crop from any reference image, via SAM3 text-prompted segmentation.
The controls are the required face_id_enabled toggle plus a small cluster of identity knobs:
face_id_enabled- master switch. Default off; V3 behaves like V2 when it's off.identity_projector/source_id/arcface_mode- which face-embedding projector pipeline to use and how it reads the reference. Defaults are the sane starting point; changing these is for people who've already seen the defaults fail.phase_scaleandid_strength- how hard and at what stage of sampling the identity is pushed.id_strengthis your dial: too low and identity drifts anyway, too high and the face starts looking pasted-on and stiff.
The rest of the panel
Since V3 inherits V2, you get the full director: the timeline editor with MAIN/AUDIO/BG AUDIO/MOTION tracks, [SPEECH]/[SOUNDS] tags (uppercase only), Seed Hunt with seed_hunt and the use_seed_hunt_1..4 toggles, chunking (chunk_duration_seconds, carry_frames), and the two-stage sampling (stage1_steps 8, stage2_steps 4, stage2_denoise 0.42, cfg 1). Outputs are last_chunk_frames, audio, stage1_frames, and the four seed-hunt preview/audio pairs.
Installing it
Same pack, same drill:
cd ComfyUI/custom_nodes
git clone https://github.com/muse-collective-26/muse-ltx-timeline
Restart and pip install av torchaudio soundfile. You'll need the LTX 2.3 stack plus the talking-head LoRA if you're doing lipsync. If you want Face ID plus the automatic face crop, add MuseFaceLock - which needs the comfyui_sam3 custom node package, a dependency the README doesn't advertise. The pack loads V3 as a WIP module in a try/except, so a missing dependency won't take the whole pack down, but it will silently skip the node.
Gotchas
- Face ID is not magic. A close-up reference with good lighting gives you a chance; a mid-shot of a face at an angle gives you a coin flip. Crop tight, crop well.
- Identity conditioning costs VRAM and time on top of an already-heavy 22B pipeline.
- V3's Face ID sits alongside V4's Ghost Mask as two different answers to the same problem - V3 locks the face, V4 locks whole-character references across the timeline. If you're fighting full-character drift rather than just faces, V4 is the one.
It's WIP, it's fiddly, and the identity knobs reward tuning. But if you've ever watched a character's face quietly change over a ninety-second clip, you know why this node exists.
Inputs (72)
| Name | Type | Default | Description |
|---|---|---|---|
| model | MODEL | — | |
| clip | CLIP | — | |
| audio_vae | VAE | — | |
| vae | VAE | — | |
| spatial_upscaler | LATENT_UPSCALE_MODEL | — | |
| start_second | FLOAT | 0.000–3600 | — |
| end_second | FLOAT | 10.000–3600 | — |
| duration_seconds | FLOAT | 10.000–3600 | — |
| start_frame | INT | 00–86400 | — |
| end_frame | INT | 2400–86400 | — |
| duration_frames | INT | 2401–86400 | — |
| timeline_data | STRING | {} | — |
| local_prompts | STRING | — | |
| segment_lengths | STRING | — | |
| global_prompt | STRING | — | |
| guide_strength | STRING | — | |
| epsilon | FLOAT | 0.00100–1 | — |
| frame_rate | FLOAT | 24.001–120 | — |
| display_mode | COMBO | seconds | 2 options: seconds, frames |
| custom_width | INT | 96064–4096 | — |
| custom_height | INT | 54464–4096 | — |
| resize_method | COMBO | maintain aspect ratio | 4 options: maintain aspect ratio, stretch to fit, crop, pad |
| divisible_by | INT | 321–256 | — |
| img_compression | INT | 180–51 | — |
| generate_audio | BOOLEAN | true | LTX generates ambient/sfx audio from [SOUNDS] prompts. |
| custom_audio_on | BOOLEAN | false | Use audio file(s) from the AUDIO timeline track. |
| lipsync | BOOLEAN | true | Sync mouth movements to custom audio. Requires Custom Audio ON and talking head LoRA. |
| motion_guide_on | BOOLEAN | true | Use motion guide segments from the timeline. |
| chunk_duration_seconds | FLOAT | 10.02–120 | — |
| auto_chunk_threshold | FLOAT | 10.00–3600 | — |
| carry_frames | INT | 731–240 | Reference frames from previous chunk locked at chunk start. 73 ≈ 3s at 24fps. |
| carry_strength | FLOAT | 1.000–1 | — |
| crossfade_frames | INT | 00–120 | — |
| ic_lora_name | COMBO | None | 1 options: None |
| ic_lora_strength | FLOAT | 1.00-10–10 | — |
| stage1_steps | INT | 81–50 | — |
| stage2_steps | INT | 41–50 | — |
| stage2_denoise | FLOAT | 0.420–1 | — |
| cfg | FLOAT | 1.00–20 | — |
| seed | INT | 420–18446744073709550000 | — |
| filename_prefix | STRING | muse | — |
| bg_volume | FLOAT | 1.000–2 | — |
| guide_scale_by | FLOAT | 0.500.01–8 | — |
| guide_scale_by_s2 | FLOAT | 1.000.01–8 | — |
| guide_upscale_method | COMBO | bicubic | 5 options: bicubic, bilinear, nearest-exact, area, bislerp |
| guide_image_attn_strength | FLOAT | 1.000–1 | — |
| guide_crop | COMBO | center | 2 options: center, disabled |
| guide_auto_snap_ic_grid | BOOLEAN | true | — |
| guide_use_tiled_encode | BOOLEAN | false | — |
| guide_tile_size | INT | 25664–512 | — |
| guide_tile_overlap | INT | 6416–256 | — |
| timeline_ui | STRING | — | |
| seed_hunt | BOOLEAN | false | ON + no candidate chosen: run a 4-seed Stage-1-resolution preview instead of the full pipeline. ON + one use_seed_hunt_N chosen: commit to that candidate — Stage 2 refines its actual cached latent instead of regenerating Stage 1 from scratch. |
| seed_hunt_steps | INT | 61–50 | — |
| seed_hunt_scale | FLOAT | 0.250.05–1 | Unused as of 1.0.4 — Seed Hunt now scouts at Stage 1's real resolution automatically (so the picked candidate's actual latent can carry forward into Stage 2). Kept as a widget only so older saved workflows still load correctly. |
| seed_hunt_1 | INT | 10–18446744073709550000 | Unused as of 1.0.4 — scouting now draws a fresh random seed for each candidate every run instead of reusing these fixed values (the actual latent carries forward on commit, so the seed number no longer needs to be fixed or reproducible). |
| seed_hunt_2 | INT | 20–18446744073709550000 | — |
| seed_hunt_3 | INT | 30–18446744073709550000 | — |
| seed_hunt_4 | INT | 40–18446744073709550000 | — |
| use_seed_hunt_1 | BOOLEAN | false | — |
| use_seed_hunt_2 | BOOLEAN | false | — |
| use_seed_hunt_3 | BOOLEAN | false | — |
| use_seed_hunt_4 | BOOLEAN | false | — |
| face_id_enabled | BOOLEAN | false | Patches the model at every Stage 1/Stage 2/per-chunk build point so the sampled face is pulled toward face_reference_image. Requires the matching Best-Face-ID LoRA already loaded onto the model input, and ComfyUI-BFSNodes installed. Off = identical to V2. |
| identity_projector | STRING | None | ArcFace projector .safetensors filename from models/loras, or 'None' for overlap-only (recommended default — the projector is a weak channel; the overlap latent carries the bulk of identity). |
| source_id | FLOAT | 20–8 | — |
| phase_scale | FLOAT | 1.00–4 | — |
| id_strength | FLOAT | 1.00–50 | — |
| arcface_mode | COMBO | auto_adjust | 3 options: auto_adjust, as_is, disable |
| bg_audioopt | AUDIO | — | |
| base_modelopt | MODEL | Base model without talking-head LoRA. Connect the UNETLoader output directly here so the ambient audio pass generates sounds without speech. | |
| face_reference_imageopt | IMAGE | Reference face for Face ID (see face_id_enabled). Feed it a close-up face crop — Muse Face Lock can produce one from any reference image automatically. |
Outputs (11)
| Name | Type | Description |
|---|---|---|
| last_chunk_frames | IMAGE | — |
| audio | AUDIO | — |
| stage1_frames | IMAGE | — |
| seed_hunt_preview_1 | IMAGE | — |
| seed_hunt_preview_2 | IMAGE | — |
| seed_hunt_preview_3 | IMAGE | — |
| seed_hunt_preview_4 | IMAGE | — |
| seed_hunt_audio_1 | AUDIO | — |
| seed_hunt_audio_2 | AUDIO | — |
| seed_hunt_audio_3 | AUDIO | — |
| seed_hunt_audio_4 | AUDIO | — |