WanAnimate2 To Video (自定义)
The official animate node, with your reference images still working
- positive
- negative
- vae
- clip_vision_output
- reference_image
- pose_video
- positive_pose
- clip_vision_output_pose
- continue_motion
- positive
- negative
- latent
- trim_latent
- trim_image
- video_frame_offset
- concat_latent
What it is, and why it exists
ComfyUI ships WanAnimate2ToVideo as a core node, and it works: it takes the Wan 2.2 Animate model (character animation and replacement, built on the Wan 2.1 I2V base) and injects the pose branch as key/value rather than piling everything into one concat. Two things are missing. It won't take a batch of reference images - the six-angles-of-your-face trick first-gen WanAnimate users rely on - and the extension path has a reputation: the classic report is that chunk two replays the first chunk with a slow zoom (1nwrlgx, "WanAnimate Comfy native does not extend").
This pack's node is a fusion rather than a wrapper. You get the second-gen mechanism (the pose branch, plus pose_strength, the pose time window, reference_image_strength, a pose-specific prompt and CLIP vision) and the first-gen node's habits on top: reference images encoded one at a time, the fix/legacy/vanilla mask modes, and a concat_latent output a context-window node can chew on. Wan 2.2 is the last open Wan and SCAIL-2 now owns the motion-transfer default, so you're here for the pose path specifically.
How it actually works
The node VAE-encodes each reference image separately - one image, one latent frame - and concatenates them on the time axis. That per-image encoding is deliberate: it's what keeps EverAnimate-style LoRAs happy, and the old "1+4n batch encode" was dropped because it behaved badly here. Then it builds the image zone: a full-length grey placeholder, filled at the head with your continuation frames if you supplied any. The two blocks become the concat latent the model consumes, and a single-channel concat mask marks the reference/continuation region as known (0) and the rest as to-generate (1).
The pose branch is where the interesting fix lives. Your pose_video gets upscaled to the output resolution, VAE-encoded, then zero-padded at the front by trim_latent - 1 frames, because the second-gen model asserts pose_latents.shape[2] == f_gen - 1. Without that padding, multi-reference runs fail the check.
The inputs you'll actually touch
reference_image- the raw batch (e.g.selected_imagesfrom the pack's Reference Image Selector). Do not pre-splice the 1+4n pattern yourself; this node encodes per image on its own.pose_video- your skeleton render as an IMAGE sequence. Sliced fromvideo_frame_offsetonward, padded by repeating the last frame if it runs short.length(default 81, step 4) andwidth/height(multiples of 16). Shavelengthfirst if you're tight on VRAM.pose_strength(1.0 = trained behaviour) and the pose time window,pose_start_percent/pose_end_percent. Motion is largely established early, sopose_end_percent = 0.7loosens fine detail while keeping the choreography - and it's faster, because the pose branch is skipped outside the window.reference_image_strength- below 1 loosens identity adherence, above 1 tightens it against drift. Reach for this before you blame the LoRA.- For chunked runs:
continue_motion,continue_motion_max_frames(5 is sane) andvideo_frame_offsetfrom the previous chunk's output of the same name.
mode defaults to vanilla; fix and legacy (plus tail_frame_count, the tail strengths, mid_frame, neutral_mix_min/max) handle black frames carried over from a previous chunk. Leave them alone until you see black frames in your continuation input.
Outputs
positive and negative carry the injected concat latents into your KSampler, and latent is the empty latent shaped to the total frame count. concat_latent is the full strip - decode that one. trim_latent / trim_image tell downstream nodes how many head frames are reference/continuation rather than generated output, so use them instead of eyeballing a crop. video_frame_offset comes back incremented by length: wire it into the next chunk, or every chunk re-predicts from frame one.
Installing it
cd ComfyUI/custom_nodes
git clone https://github.com/user2318/ComfyUI-CustomNodeKit.git
cd ComfyUI-CustomNodeKit
pip install -r requirements.txt
Restart ComfyUI, or search "ComfyUI-CustomNodeKit" in ComfyUI Manager. That's the whole install for this node - the pack's install.py also pip-installs groundingdino-py and transformers for its unrelated detection nodes on first import. If that part fails, the pack still loads.
No models ship with it: you still need the Wan 2.2 Animate UNet, the Wan VAE, CLIP vision and a text encoder.
Where people get burned
pose_start_percent > pose_end_percent isn't a sampler crash, it's an immediate ValueError from the node. It reads like a framework bug the first time.
If a chunk comes out looking like the previous one with a slow zoom, check the wiring first. video_frame_offset has to actually reach this node, and pose_video has to be long enough to slice from that offset - the node drops everything before it, then pads by repeating the final frame, so a stale offset gives you a plausible-looking run with the wrong skeleton throughout. That's the honest answer to most "the extend nodes don't extend" complaints against the native workflow.
One upgrade gotcha: v1.8.0 removed WanSCAILSparseAttention outright - O(S²) attention masks that exploded memory on long sequences for no measurable benefit. Old workflows will just show a missing node; delete the orphan.
Inputs (28)
| Name | Type | Default | Description |
|---|---|---|---|
| positive | CONDITIONING | 正向提示词 conditioning。Positive prompt conditioning. | |
| negative | CONDITIONING | 负向提示词 conditioning。Negative prompt conditioning. | |
| vae | VAE | Wan 模型的 VAE,用于编码参考图和窗口帧。VAE for the Wan model, used to encode reference images and window frames. | |
| width | INT | 83216–16384 | 生成视频的宽度(像素),必须是 16 的倍数。Width of the generated video (pixels), must be a multiple of 16. |
| height | INT | 48016–16384 | 生成视频的高度(像素),必须是 16 的倍数。Height of the generated video (pixels), must be a multiple of 16. |
| length | INT | 811–16384 | 生成视频的帧数(像素帧),必须是 4 的倍数。Number of frames for the generated video (pixel frames), must be a multiple of 4. |
| batch_size | INT | 11–4096 | 批量大小,通常保持为 1。Batch size, usually keep at 1. |
| continue_motion_max_frames | INT | 51–16384 | 从上一块携带的最大帧数(RGB 图像),用于块间接续。Maximum frames carried from the previous chunk (RGB images), used for inter-chunk continuity. |
| video_frame_offset | INT | 00–16384 | 当前 chunk 的帧偏移量,从上一块的 video_frame_offset 输出接入。Frame offset of the current chunk, connected from the previous chunk's video_frame_offset output. |
| transition_width | INT | 00–128 | fix 模式下黑帧区域的过渡区宽度,0=禁用过渡。Transition zone width for black frame areas in fix mode, 0=disable transition. |
| mode | COMBO | vanilla | 掩码模式:vanilla=官方行为(fix/legacy不动),fix=黑帧检测+过渡,legacy=尾帧渐变。Mask mode: vanilla=official behavior, fix=black frame detection+transition, legacy=tail frame fade. |
| tail_frame_count | INT | 00–1000 | legacy 模式下尾帧处理帧数,0=禁用。Number of tail frames to process in legacy mode, 0=disabled. |
| tail_start_strength | FLOAT | 0.000–1 | legacy 模式尾帧起始强度。Tail frame start strength in legacy mode. |
| tail_end_strength | FLOAT | 0.000–1 | legacy 模式尾帧结束强度。Tail frame end strength in legacy mode. |
| clip_vision_outputopt | CLIP_VISION_OUTPUT | CLIP Vision 输出,用于参考图语义理解。CLIP Vision output for reference image semantic understanding. | |
| reference_imageopt | IMAGE | 参考图输入。接参考图选择器的 selected_images(排序后的原始图片,节点内部自动逐帧编码)。Reference image input. Connect to selected_images from ReferenceImageSelector (node encodes each frame independently). | |
| pose_videoopt | IMAGE | 姿态视频帧序列,用于姿态引导。Pose video frame sequence for pose guidance. | |
| positive_poseopt | CONDITIONING | 姿态分支专属提示词(描述动作而非角色),默认回退到 positive。Prompt for the pose-video branch, describing the motion rather than the character. Defaults to positive. | |
| clip_vision_output_poseopt | CLIP_VISION_OUTPUT | 姿态视频首帧的 CLIP Vision 输出,默认回退到 clip_vision_output。CLIP vision of the pose video's first frame. Defaults to clip_vision_output. | |
| continue_motionopt | IMAGE | 上一块的末尾 RGB 帧,用于块间运动接续。Last RGB frames of the previous chunk for inter-chunk motion continuity. | |
| mid_frameopt | INT | -1-1–1000 | legacy 模式中间帧锚点位置,-1=不使用。Mid-frame anchor position in legacy mode, -1=not used. |
| mid_strengthopt | FLOAT | 0.500–1 | legacy 模式中间帧锚点强度。Mid-frame anchor strength in legacy mode. |
| neutral_mix_minopt | FLOAT | 0.000–1 | legacy 模式掩码=0时的中性灰混合比例。Neutral gray mix ratio when mask=0 in legacy mode. |
| neutral_mix_maxopt | FLOAT | 1.000–1 | legacy 模式掩码=1时的中性灰混合比例。Neutral gray mix ratio when mask=1 in legacy mode. |
| pose_strengthopt | FLOAT | 1.000–10 | 缩放姿态视频对运动的影响强度。1.0=训练行为;0.0=静音但不会完全移除;>1=增强。Scales the pose video's influence on the motion. 1.0 is the trained behavior; 0.0 mutes it but does not fully remove it; above amplifies. |
| pose_start_percentopt | FLOAT | 0.000–1 | 姿态影响的起始采样百分比,窗口外整个姿态分支被跳过(可提速)。Sampling percent at which the pose influence starts. Outside the window the pose branch is skipped entirely. |
| pose_end_percentopt | FLOAT | 1.000–1 | 姿态影响的结束采样百分比,例如 0.7 可在保留舞蹈编排的同时放松细节。Sampling percent at which the pose influence ends. Motion is mostly established early, so e.g. 0.7 can loosen fine detail while keeping the choreography. |
| reference_image_strengthopt | FLOAT | 1.000–10 | 生成帧对参考图 latent 帧的注意力强度。<1 放松身份/外观一致性,>1 收紧。Scales how strongly generated frames attend to the reference image's latent frame. Below 1.0 loosens identity/appearance adherence, above tightens it against drift. |
Outputs (7)
| Name | Type | Description |
|---|---|---|
| positive | CONDITIONING | — |
| negative | CONDITIONING | — |
| latent | LATENT | — |
| trim_latent | INT | — |
| trim_image | INT | — |
| video_frame_offset | INT | — |
| concat_latent | LATENT | — |