Nodes/ComfyUI-WanAnimatePlus/WanAnimatePlus Animate2 Embeds
ComfyUI Node

WanAnimatePlus Animate2 Embeds

Official Animate 2 conditioning, with room for extra reference faces

By wuwukaka·Created 4 months ago·Updated 14 days ago· 415
WanAnimatePlus Animate2 Embeds
  • positive
  • negative
  • vae
  • clip_vision_output
  • clip_vision_output_pose
  • positive_pose
  • ref_image
  • bg_image
  • pose_images
  • prefix_frames
  • positive
  • negative
  • latent
◄width832►
◄height480►
◄num_frames81►
◄frame_window_size81►
◄batch_size1►
◄pose_strength1.00►
◄ref_strength1.00►
◄tiled_vaefalse►

What this node is for

Wan Animate gets recommended for exactly one thing: making a photo of a person do what a video of a person does, without the plastic stretch older approaches gave you. The pitch for SCAIL-2 was essentially "this, but without the stick figure" - and it landed.

There are two ways into Animate 2 in ComfyUI. Core ships WanAnimate2ToVideo (marked experimental), and kijai's WanVideoWrapper has its own chain with WANVID* types all over it. WanAnimatePlus Animate2 Embeds is the official-type conditioning node - CONDITIONING, VAE, LATENT, no wrapper objects - with three things the core node doesn't give you: prefix_frames, bg_image, and an internal-loop handoff the pack's sampler finishes. You reach for it when your graph is core types end to end (UNETLoader or a GGUF loader, VAELoader, CLIPTextEncode) and you'd rather not rebuild the pipeline in wrapper dialect.

Be honest about the pack's standing, though: it's a one-person fork of kijai's WanVideoWrapper by wuwukasi (Apache 2.0), and the Animate2 nodes are its newest commit - the README doesn't document them yet, so this comes from reading the source. The only community signal I found for the repo is someone recommending it as the way to get motion transfer with subject reference images. You're early.

How it works

Animate 2 conditioning is a latent concat plus a mask. That's the whole trick. ref_image is VAE-encoded into a single identity latent frame; a canvas of num_frames pixel frames is encoded; the identity latent is concatenated in front of it on the time axis. A mask marks those front frames as known (0, frozen) and the rest free (1), so the reference stays pinned while everything else denoises. pose_images become a pose_video_latent with clip-vision and cross-attention from positive_pose, so the motion branch actually drives it.

Everything this node adds sits on that base. bg_image fills the canvas with your plate instead of the mid-grey the core node writes - if you're compositing a subject onto a background, that grey is what you spend an afternoon fighting. prefix_frames takes up to five more images and encodes each at full resolution as its own identity latent in the reference stream: more face angles, the other sleeve. The README suggests three as the practical number.

frame_window_size is the interesting one. Leave it equal to num_frames (both default to 81) and the node builds the full conditioning now, then releases the VAE. Set it below num_frames and it deliberately builds nothing - it hands the sampler an empty latent plus a runtime record, and the sampler runs the chunked loop with a 5-frame handoff. Same maths SCAIL-2 Infinity uses: 81-frame windows, stepping 76.

Sizes get snapped for you: width/height down to multiples of 32, frame counts down to 4n+1. Ask for 100 frames, get 97 - the Wan VAE compresses time roughly 4x, and that's why your clip is always a couple of frames short.

The inputs that matter

Required: positive, negative, vae, width/height (832x480 default), num_frames (81), frame_window_size (81 - lower it for long clips), batch_size, pose_strength, ref_strength.

Those last two are the dials worth understanding. 1.0 is trained behaviour. Below 1 loosens adherence, handy when a prompt needs to restyle the subject. Above 1 tightens it, which is what you reach for when identity drifts over a long take. The node only writes them into the conditioning when they aren't 1.0.

Optional: ref_image, bg_image, pose_images, prefix_frames, tiled_vae (turn it on if VAE encoding is what's blowing up VRAM), clip_vision_output, clip_vision_output_pose, positive_pose.

Outputs are positive, negative - your conditioning, rewritten with the Animate 2 values - and latent. All three go to WanAnimatePlus Animate2 Sampler.

Install

ComfyUI Manager, search WanAnimatePlus. Or:

cd ComfyUI/custom_nodes
git clone https://github.com/wuwukaka/ComfyUI-WanAnimatePlus.git

Restart ComfyUI. requirements.txt is the usual video-model pile - diffusers>=0.33.0, peft>=0.17.0, accelerate, gguf>=0.17.1, opencv-python, scipy, pyloudnorm and friends. The pack's FAQ also tells you to keep the original ComfyUI-WanVideoWrapper installed alongside it.

For models, load Animate 2 through ComfyUI's own loaders: UNETLoader (or a GGUF loader), VAELoader, CLIPLoader. Don't use the pack's WanAnimatePlus ModelLoader here - it outputs the wrapper's WANVIDEOMODEL type, which these nodes can't eat.

Where people get burned

Mixing dialects mid-graph. The README's own warning: mixing WanAnimatePlus nodes with WanVideoWrapper nodes in one workflow degrades output. With these two, the chain is Animate2 Embeds → Animate2 Sampler → WanAnimatePlus VAE Decode. Keep it that way.

Nodes not showing up. Nine times in ten the folder isn't named ComfyUI-WanAnimatePlus, or you didn't restart.

The VAE has to reach the sampler too. In one-shot mode the embeds releases it after building conditioning, so wire the same VAE into the sampler or it will stop you with an error about needing a VAE to trim identity latents.

The 81-frame ceiling is real. Long Wan Animate takes have always degraded past the first batch; the old answer was manual chunking, and now it's nodes like this one. Lower frame_window_size and let the sampler loop - and expect some identity drift at the seams anyway, because that's the model, not your settings.

CategoryWanAnimatePlus

Inputs (18)

NameTypeDefaultDescription
positiveCONDITIONING—
negativeCONDITIONING—
vaeVAE—
widthINT83264–8096—
heightINT48064–8096—
num_framesINT811–10000—
frame_window_sizeINT811–10000—
batch_sizeINT11–4096—
pose_strengthFLOAT1.000–10—
ref_strengthFLOAT1.000–10—
clip_vision_outputoptCLIP_VISION_OUTPUT—
clip_vision_output_poseoptCLIP_VISION_OUTPUT—
positive_poseoptCONDITIONING—
ref_imageoptIMAGE—
bg_imageoptIMAGE—
pose_imagesoptIMAGE—
prefix_framesoptIMAGE—
tiled_vaeoptBOOLEANfalse—

Outputs (3)

NameTypeDescription
positiveCONDITIONING—
negativeCONDITIONING—
latentLATENT—