Video Temporal Upscale
Upscale a clip without the wallpaper crawling
- source
- upscale_model
- motion_model
- ema_vfi_model
- video
- frames
Everyone finds this out the same way. You upscale a video frame by frame with a 4x model you've used happily on stills for two years, and fine repeating texture - patterned wallpaper, gravel, foliage - starts crawling. Not blurry, not wrong: jittering, because the model made slightly different guesses on frame 41 than on frame 40, and nothing in a per-frame upscaler forces consistency.
Video Temporal Upscale fixes that without changing models: it's in the same family as the pack's EMA-VFI interpolation and motion nodes - take the detail the upscaler invented and hold it steady along the clip's own motion.
How it works
Here's the trick. The node runs your upscale model over each frame - tiled, so VRAM doesn't have to fit the whole frame - but it doesn't keep the upscale. It computes a plain Lanczos resize of the source, subtracts that from the model's output, and keeps the difference: the residual, literally the detail the model made up.
That residual is then carried along the clip's measured optical flow, backwards and forwards, blended with a neighbour's residual only where the pixels look like the same surface - a luminance gap over about 12/255 reads as "different thing" or a cut, and stops the carry there. Finally the steadied residual goes back onto the plain resize, so the invented detail is one continuous thing along the motion instead of a fresh decision per frame.
strength is the dial on that: 0 upscales each frame on its own (fast, flickery, the old behaviour), 0.8 is the default, 1.0 holds the detail until a surface actually changes.
The inputs that matter
- source - a VIDEO or IMAGE frames. A VIDEO brings its frame rate and audio; bare frames play at 24fps and are silent.
- upscale_model - from Load Upscale Model. A 4x model can produce a 2x result;
upscale_factoris the final size against the source, rounded to even pixel sides. - strength - how much of each frame's detail comes from its neighbours. Start at the 0.8 default.
- tile_size and overlap - 512 and 32 suit most cards. If the card runs out, the tile is halved and the frame retried automatically; set 256 yourself if you'd rather not watch it figure that out.
- frame_rate - 0 keeps the source's; 30 gives 30fps. The length is kept, so audio stays in sync, and frames that fall between source frames are drawn by EMA-VFI if you wire
ema_vfi_model(from the pack's EMA-VFI Video Model Loader). Without it,new_framesfalls back tohold- the nearer frame repeats, which looks exactly as good as it sounds. Note the pack's own warning: any rate but a straight double needs anours_tcheckpoint. - crf - 16 by default, x264 quality of the written file.
- Optional motion_model (SEA-RAFT or FlowSeek, from Video Motion Model Loader) replaces the built-in texture-flow estimate, and precision picks
autoor32 bit float.
Outputs are video, the upscaled clip as an h264 file with the source's audio, and frames, an INT. It's an output node - it writes the file as it works, which is why a 4x clip that never fits in memory still finishes. Wire the VIDEO into Save Video (Advanced) and leave its codec on ComfyUI Auto: that copies the file through instead of encoding it a second time.
Install
ComfyUI Manager, search WAS Node Suite v3. Or:
cd ComfyUI/custom_nodes
git clone https://github.com/WASasquatch/was-node-suite-comfyui.git
ComfyUI 0.14.0+ and Python 3.10+; restart. Nothing is pip-installed, now or at startup - that's the v3 behaviour, and a change from the old WAS Node Suite you may remember as dependency-brittle. The EMA-VFI and motion weights are not bundled: with features.network off in the pack's config.yaml, put checkpoint files in the folders the loaders name and restart so they appear in the dropdown.
Common issues
Out of memory, twice over. The upscale halves its tile and retries, so it mostly copes. The residuals - the detail carried between passes - are a separate budget, and the node raises MemoryError when neither free RAM nor a scratch drive can hold them. Keep a few GB free on the ComfyUI temp drive before a long 4x pass.
HDR clips are refused. The node writes sRGB and raises ValueError if the clip isn't in an sRGB space. Load an sRGB version; there's no "just do it anyway" toggle.
Interpolation without a checkpoint silently degrades. Leave ema_vfi_model empty and new_frames falls back to holding the nearer frame - judder that looks like a bug in the upscale. If the result stutters after you changed the frame rate, that's your cause.
The comparison is right there on the node. The player shows your clip both ways - the plain per-frame upscale and the steadied result - so you can judge whether the temporal pass earned its runtime on your footage. On a static shot it changes almost nothing. On a pan across texture, it's the difference between watchable and not.
Inputs (12)
| Name | Type | Default | Description |
|---|---|---|---|
| source | VIDEO,IMAGE | The clip to upscale, as a VIDEO or IMAGE frames. A VIDEO gives the result its frame rate and audio; frames play at 24 fps with none. | |
| upscale_model | UPSCALE_MODEL | The upscale model, from Load Upscale Model. A 4x model can make a 2x result. | |
| upscale_factor | FLOAT | 4.01–16 | Final size against the source: 2.0 = double; 4.0 = four times. Sides are rounded to even numbers of pixels. |
| strength | FLOAT | 0.800–1 | How much of each frame's detail comes from its neighbours: 0 = each frame upscaled on its own; 0.8 = default; 1.0 = holds the detail until a surface changes, or on a clip too large to hold at once, until the next block of frames. |
| frame_rate | FLOAT | 0.0000–240 | Frames per second of the result: 0 = the source's; 30 = 30 fps. The length is kept, so the audio stays in sync. |
| tile_size | INT | 51264–4096 | Tile edge in source pixels: 256 and 512 suit most cards. If the card runs out, the tile is halved and the frame retried. |
| overlap | INT | 320–1024 | How far neighbouring tiles overlap, in source pixels: 0 = hard joins; 32 to 64 hides them on most models. |
| crf | FLOAT | 16.00–51 | Quality of the written clip, lower is better and larger: 0 = lossless; 16 = default; 23 = typical; 28 = small. |
| motion_modelopt | WAS_MOTION_MODEL | A flow network from Video Motion Model Loader, which the motion is measured with. Left empty, the built-in texture flow measures it. | |
| ema_vfi_modelopt | EMA_VFI_MODEL | The interpolation network from EMA-VFI Video Model Loader, which draws frames between two when frame_rate changes. Left empty, the nearer frame repeats. Any rate but double needs an 'ours_t' checkpoint. | |
| precisionopt | COMBO | auto | What the upscale model runs in: 'auto' = half precision where the model declares it safe; '32 bit float' = every model's safest. |
| new_framesopt | COMBO | auto | How a frame between two source frames is made when frame_rate changes: 'auto' = EMA-VFI where the motion accounts for the change, the nearer frame where it does not, such as a mouth changing shape; 'interpolate' = EMA-VFI always; 'blend' = the two mixed; 'hold' = the nearer frame. |
Outputs (2)
| Name | Type | Description |
|---|---|---|
| video | VIDEO | The upscaled clip as an h264 file, with the source's audio. Save Video on 'auto' keeps it without encoding it again. |
| frames | INT | How many frames the upscaled clip holds. |