Context Window Anchor (窗口首帧锚定)
Pinning every Wan window to frame zero (and why the author says test it first)
- model
- global_pose_frames
- global_mask_frames
- vae
- model
The problem it's aimed at
Wan's native context is 81 frames, so anything longer is windowing: slice the sequence into overlapping windows, sample each against the same model, fuse the results with a weight schedule. That works, but it has a specific failure. Animate2 and SCAIL-2 derive their "reference image ↔ driving video" body-scale and position alignment per window, from whatever driving frames that window contains. Window one starts on one pose, window four on another, so each window anchors itself differently - and the fusion shows it as a jump at the seam.
Long-context runs generate more error than answer in general - window-size mismatches that still error out, and Expected all tensors to be on the same device blow-ups the moment you plug reference latents into a context setup (1pao1t6, 1pc8b7z). And the author labels this node experimental and test-only: effectiveness unproven, side effects on quality, identity and sampling stability unevaluated.
How it works
It edits conditioning, at the point where the context handler re-slices each window - no pixels, no latents. The required model input has to be the output of a context-window node (Custom Context Windows (Manual), or Wan SCAIL-2 Context Windows for SCAIL-2), because those are what stash a handler on model.model_options["context_handler"]. This node clones the model, shallow-copies that handler and wraps its get_resized_cond, so after the per-window slice, windows after the first get their leading pose_latents frames overwritten with window 0's (and on SCAIL-2, the driving mask sam_latents too).
Three consequences fall out of doing it there rather than on the latent:
- Every sampling step, because
get_resized_condruns every step. - Never changes tensor lengths, so animate2's
pose_latents.shape[2] == f_gen - 1check still passes. - Never crosses chunks: the anchor source is window 0 of the current segment, so drift between segments is untouched. It's a within-segment seam fix, and the author says so.
The anchor_offset default of -1 is the clever bit. Context nodes prepend the prefix reference frames to each window's index list, and the driving-video slice comes from that same list - so every window's slice starts with P-1 prefix-derived frames that are identical across windows, and anchoring there would do nothing. Auto offset skips them: six reference images → offset 5.
The inputs worth setting
anchor_frames(default 1, max 32, 0 disables) - start at 1, try 2–3 after that. Turn onlog_debugfor one run: it prints every window hit with shapes and first-frame values.anchor_targets- "驱动视频 + 驱动遮罩 (auto)", "pose only" or "mask only". On SCAIL-2, cycling through these shows which of the two is dragging your seam around.anchor_stepsvsblend- two strategies, not two ends of one slider.blendbelow 1.0 mixes the anchor with the window's own frames, feeding a synthetic in-between every step.anchor_steps = Nanchors for the first N steps, then returns to real conditioning. Tryblend=1.0 + anchor_steps=1againstblend=0.7 + anchor_steps=-1.global_pose_frames(plusglobal_mask_framesandvaeon SCAIL-2) - hybrid anchoring. Feed it the head of the whole driving video, taken before your loop trims the clip to the current segment. More frames than the anchor band (M = ((T-1)//4)+1) means full replacement, fewer means partial. Optional - unconnected, it behaves exactly like plain per-segment anchoring.
Output is a single model: drop it between the context node and the KSampler.
Installing it
Same pack, same three steps:
cd ComfyUI/custom_nodes
git clone https://github.com/user2318/ComfyUI-CustomNodeKit.git
cd ComfyUI-CustomNodeKit
pip install -r requirements.txt
Restart, or search "ComfyUI-CustomNodeKit" in ComfyUI Manager. No model files, no extra dependencies - it's pure Python against the model object. (The pack's install.py does pull groundingdino-py and transformers for its unrelated detection nodes; if that fails, the rest loads anyway.)
Where people get burned
The node does nothing and doesn't complain. If the model doesn't carry a context_handler, it logs a warning and passes the model through - no error, no red node, and you conclude the setting didn't help. Usual cause: it's wired before the context node instead of after, or a Fast Muter upstream cut that node out of the path.
No change with log_debug on. Only tensors split per window get touched; whole-sequence conditions are skipped deliberately, and SCAIL's reference-frame tensors are in that category. Also check enabled, anchor_frames > 0 and anchor_steps != 0 - all three pass through when off.
blend won't soften the mask. On mask keys the blend is coerced to 1.0 by design: sam_latents is category data (colour coverage per identity), and mixing two frames of it produces boundaries that match neither. If the mask anchor causes jitter, use anchor_steps to switch it off early instead. A shape-mismatch warning means source and target shapes disagree and that key was skipped.
And the caveat that decides whether to use it at all: same seed, one variable at a time, node in versus node out, twice. If your seams already look clean, leave it out of the graph.
Inputs (12)
| Name | Type | Default | Description |
|---|---|---|---|
| model | MODEL | 已接入上下文窗口节点的模型(Custom Context Windows (Manual) 或 Wan SCAIL-2 Context Windows 的输出)。Model already wrapped by a context windows node. | |
| enabledopt | BOOLEAN | true | 是否启用锚定;关闭时模型直接透传。Whether to enable anchoring; when off the model is passed through unchanged. |
| anchor_targetsopt | COMBO | 驱动视频 + 驱动遮罩 (auto) | 要锚定的目标(下拉选择)。驱动视频=pose 分支(animate2 / SCAIL-2 都是 pose_latents);驱动遮罩=SCAIL-2 的 driving_mask(sam_latents)。Targets to anchor: the driving video (pose branch) and/or the SCAIL-2 driving mask. |
| anchor_framesopt | INT | 10–32 | 每个窗口要锚定的 latent 帧数,0=关闭。建议先 1,再试 2-3。Number of leading latent frames to anchor per window; 0 disables. |
| anchor_offsetopt | INT | -1-1–64 | -1=自动:跳过切片头部的“前缀派生帧”(=max(0, 前缀帧数-1)),从窗口自身的第一个驱动帧开始锚定;填 0/1/… 可手动指定起始帧做实验。-1 = auto: skip prefix-derived frames at the slice head. |
| blendopt | FLOAT | 1.000–1 | 1.0=完全替换为窗口 0 的首帧;<1.0 与窗口自身首帧线性混合(用于缓解首帧姿态回跳造成的抖动)。1.0 fully replaces with the anchor; below 1.0 linearly blends with the window's own frames. |
| apply_to_negativeopt | BOOLEAN | true | CFG 下是否同时锚定 negative(通常两者传入同一份驱动条件,锚定结果一致,不会造成 CFG 偏斜)。Also anchor the negative conditioning under CFG. |
| log_debugopt | BOOLEAN | false | 打印每个窗口命中的键、shape 与首帧数值,用于验证锚定是否生效。Log per-window hits, shapes and first-frame stats for debugging. |
| anchor_stepsopt | INT | -1-1–64 | 只在前 N 个采样步锚定:-1=全程(默认);0=不锚定;N>0=前 N 步。时间门控(每步给真实信号,而非幅度混合的合成中间态)。-1 = all steps (default); N > 0 = only the first N sampling steps. |
| global_pose_framesopt | IMAGE | 【可选】驱动/pose 视频开头的若干像素帧(给多少用多少,不设帧数参数)。编码成 latent 后替换锚定带的前 M 帧(M=((T-1)//4)+1):给多了=完全替换,给少了=部分替换;不接=纯本段锚定。Optional global driving-video head frames. | |
| global_mask_framesopt | IMAGE | 【可选,SCAIL-2】彩色遮罩视频开头的若干像素帧,按 _extract_mask_to_28ch 编码后替换遮罩锚的前 M 帧。Optional global colored-mask head frames (SCAIL-2). | |
| vaeopt | VAE | 【可选】仅用于把 global_pose_frames 编码成 latent(接同一个 Wan VAE);遮罩池不需要 VAE。VAE used to encode global_pose_frames. |
Outputs (1)
| Name | Type | Description |
|---|---|---|
| model | MODEL | — |