Video Temporal Consistency
Stop per-frame effects from flickering
- source
- processed
- motion
- images
- video
Here is the problem this node solves, and it is one of the most annoying in all of video work. You have a clip and an effect that only exists for still images: a style filter, a grade, an LUT, an upscaler, a depth model, a detailer. You run it on every frame. Individually the frames look great. Played back, the effect boils - texture crawls, edges shimmer, the grade flickers - because the effect has no idea frame two exists.
The usual workarounds are generative: re-pass the whole clip through a video model and lose the effect you carefully built. Video Temporal Consistency does it the compositor's way instead. It reads the motion from the source clip, carries what the effect changed along that motion from the frames around it, and holds it steady - so the effect sticks to the surfaces it was painted on while the clip itself moves exactly as it did. You keep the clip's motion and the effect's look, and you skip the re-render.
Inputs
source is the clip before the effect - IMAGE frames or a VIDEO. This is where the motion is read from, so it must be the original clip, not the output. If source is a VIDEO, the video output also inherits its frame rate and audio, which is a good reason to hand it a VIDEO rather than a frame batch.
processed is the same frames after the effect - as many as source, at any size. Any size is the important bit: this is a legitimate way to upsample an effect that ran at a different resolution, because the steadying happens along the motion of the source.
strength is how much of each frame's effect comes from its neighbours: 0 is the processed frames as they are (no steadying at all), 0.8 is the default, 1.0 holds the effect until a surface actually changes. Higher is more stable but sticks longer, so a fast change in the effect itself - a hard flash, an instant grade switch - gets smeared over a few frames. 0.6–0.85 is the useful band.
motion is optional and takes a shared measurement from Video Motion, measured from source or processed at any size. Wire it if your graph already has one; that is much cheaper than measuring again here at 768 px.
Outputs
images is the processed frames, steadied, at the same size and count as processed - so you can wire this straight back into whatever came next in your graph. video is the same frames as a clip, at source's frame rate and with its audio when source was a VIDEO, else processed's frame rate and audio when that was a VIDEO, otherwise 24 fps with no sound. The fallback is worth remembering: hand it two frame batches and your audio is gone.
Install
ComfyUI Manager → search WAS Node Suite v3 → install → restart. Or:
cd ComfyUI/custom_nodes
git clone https://github.com/WASasquatch/was-node-suite-comfyui.git
ComfyUI 0.14.0+ and Python 3.10+. The pack installs no packages - requirements.txt holds a comment and pip is never invoked, which is also why this node is cheap to add to an existing graph. No model weights needed; a learned flow network in ComfyUI/models/optical_flow is optional and only worth it when the built-in measurement struggles.
Where it earns its place
The obvious case is an upscaler that runs on stills. Upscale each frame, then steady it, and you avoid both the per-frame crawling and the re-generation. That is this node's best use - the upscaling doc's whole framing is that you should know which primitive does which job, and "make an image model behave consistently over time" is precisely the job a diffusion pass is bad at.
The second case is per-frame depth or segmentation: the community's own note on depth is that frame-by-frame models flicker where video-native ones don't (Depth Anything v2 vs DepthCrafter, in the depth-estimation discussions). If you cannot get a video-native model, steadying the output is the practical workaround.
Limits, stated plainly
It cannot invent consistency that isn't there. If the effect produced a genuinely different image on frame 40 - a hallucinated face, a different chair - no amount of carrying along motion fixes that; you will get a slow-morphing wrong thing instead of a flickering wrong thing.
It is not a substitute for a clip-aware model. It suppresses flicker; it does not give you a model that understands a video. For depth and style, a video-native pass will beat steadying a per-frame one.
Cuts. Motion is meaningless across a cut, so the steadying will drag a little of the previous shot into the first frames of the next. Split the clip into scenes first with Video Split Scenes if the seam is bothersome.
Higher strength is not better. At 1.0 you are holding the effect for as long as the surface is visible, and any intentional change in the effect gets dragged. Most people's best results sit around the default.
Inputs (4)
| Name | Type | Default | Description |
|---|---|---|---|
| source | IMAGE,VIDEO | The clip before the effect, as IMAGE frames or a VIDEO, which the motion is read from. A VIDEO also gives the video output its frame rate and audio. | |
| processed | IMAGE,VIDEO | The same frames after the effect, as IMAGE frames or a VIDEO, as many as source and at any size, such as an upscaler's output. | |
| strength | FLOAT | 0.800–1 | How much of each frame's effect comes from its neighbours: 0 = processed as it is; 0.8 = default; 1.0 = holds the effect until a surface changes. |
| motionopt | WAS_MOTION | Motion from Video Motion, measured from source or processed at any size. Left empty, it is measured here at 768 px on the long side. |
Outputs (2)
| Name | Type | Description |
|---|---|---|
| images | IMAGE | The processed frames, held steady: same size and count as processed. |
| video | VIDEO | The same frames as a clip, at source's frame rate and with its audio when source is a VIDEO, else processed's when that is a VIDEO, otherwise at 24 fps with none. |