Nodes/was-node-suite-comfyui/Video Temporal Upscale
ComfyUI Node Runs on cloud

Video Temporal Upscale

Upscale a clip without the wallpaper crawling

By WASasquatch·Created 4 years ago·Updated a day ago· 1,864
Video Temporal Upscale
  • source
  • upscale_model
  • motion_model
  • ema_vfi_model
  • video
  • frames
◄upscale_factor4.0►
◄strength0.80►
◄frame_rate0.000►
◄tile_size512►
◄overlap32►
◄crf16.0►
◄precisionauto►
◄new_framesauto►

Everyone finds this out the same way. You upscale a video frame by frame with a 4x model you've used happily on stills for two years, and fine repeating texture - patterned wallpaper, gravel, foliage - starts crawling. Not blurry, not wrong: jittering, because the model made slightly different guesses on frame 41 than on frame 40, and nothing in a per-frame upscaler forces consistency.

Video Temporal Upscale fixes that without changing models: it's in the same family as the pack's EMA-VFI interpolation and motion nodes - take the detail the upscaler invented and hold it steady along the clip's own motion.

How it works

Here's the trick. The node runs your upscale model over each frame - tiled, so VRAM doesn't have to fit the whole frame - but it doesn't keep the upscale. It computes a plain Lanczos resize of the source, subtracts that from the model's output, and keeps the difference: the residual, literally the detail the model made up.

That residual is then carried along the clip's measured optical flow, backwards and forwards, blended with a neighbour's residual only where the pixels look like the same surface - a luminance gap over about 12/255 reads as "different thing" or a cut, and stops the carry there. Finally the steadied residual goes back onto the plain resize, so the invented detail is one continuous thing along the motion instead of a fresh decision per frame.

strength is the dial on that: 0 upscales each frame on its own (fast, flickery, the old behaviour), 0.8 is the default, 1.0 holds the detail until a surface actually changes.

The inputs that matter

  • source - a VIDEO or IMAGE frames. A VIDEO brings its frame rate and audio; bare frames play at 24fps and are silent.
  • upscale_model - from Load Upscale Model. A 4x model can produce a 2x result; upscale_factor is the final size against the source, rounded to even pixel sides.
  • strength - how much of each frame's detail comes from its neighbours. Start at the 0.8 default.
  • tile_size and overlap - 512 and 32 suit most cards. If the card runs out, the tile is halved and the frame retried automatically; set 256 yourself if you'd rather not watch it figure that out.
  • frame_rate - 0 keeps the source's; 30 gives 30fps. The length is kept, so audio stays in sync, and frames that fall between source frames are drawn by EMA-VFI if you wire ema_vfi_model (from the pack's EMA-VFI Video Model Loader). Without it, new_frames falls back to hold - the nearer frame repeats, which looks exactly as good as it sounds. Note the pack's own warning: any rate but a straight double needs an ours_t checkpoint.
  • crf - 16 by default, x264 quality of the written file.
  • Optional motion_model (SEA-RAFT or FlowSeek, from Video Motion Model Loader) replaces the built-in texture-flow estimate, and precision picks auto or 32 bit float.

Outputs are video, the upscaled clip as an h264 file with the source's audio, and frames, an INT. It's an output node - it writes the file as it works, which is why a 4x clip that never fits in memory still finishes. Wire the VIDEO into Save Video (Advanced) and leave its codec on ComfyUI Auto: that copies the file through instead of encoding it a second time.

Install

ComfyUI Manager, search WAS Node Suite v3. Or:

cd ComfyUI/custom_nodes
git clone https://github.com/WASasquatch/was-node-suite-comfyui.git

ComfyUI 0.14.0+ and Python 3.10+; restart. Nothing is pip-installed, now or at startup - that's the v3 behaviour, and a change from the old WAS Node Suite you may remember as dependency-brittle. The EMA-VFI and motion weights are not bundled: with features.network off in the pack's config.yaml, put checkpoint files in the folders the loaders name and restart so they appear in the dropdown.

Common issues

Out of memory, twice over. The upscale halves its tile and retries, so it mostly copes. The residuals - the detail carried between passes - are a separate budget, and the node raises MemoryError when neither free RAM nor a scratch drive can hold them. Keep a few GB free on the ComfyUI temp drive before a long 4x pass.

HDR clips are refused. The node writes sRGB and raises ValueError if the clip isn't in an sRGB space. Load an sRGB version; there's no "just do it anyway" toggle.

Interpolation without a checkpoint silently degrades. Leave ema_vfi_model empty and new_frames falls back to holding the nearer frame - judder that looks like a bug in the upscale. If the result stutters after you changed the frame rate, that's your cause.

The comparison is right there on the node. The player shows your clip both ways - the plain per-frame upscale and the steadied result - so you can judge whether the temporal pass earned its runtime on your footage. On a static shot it changes almost nothing. On a pan across texture, it's the difference between watchable and not.

CategoryWAS Suite/Animation

Inputs (12)

NameTypeDefaultDescription
sourceVIDEO,IMAGEThe clip to upscale, as a VIDEO or IMAGE frames. A VIDEO gives the result its frame rate and audio; frames play at 24 fps with none.
upscale_modelUPSCALE_MODELThe upscale model, from Load Upscale Model. A 4x model can make a 2x result.
upscale_factorFLOAT4.01–16Final size against the source: 2.0 = double; 4.0 = four times. Sides are rounded to even numbers of pixels.
strengthFLOAT0.800–1How much of each frame's detail comes from its neighbours: 0 = each frame upscaled on its own; 0.8 = default; 1.0 = holds the detail until a surface changes, or on a clip too large to hold at once, until the next block of frames.
frame_rateFLOAT0.0000–240Frames per second of the result: 0 = the source's; 30 = 30 fps. The length is kept, so the audio stays in sync.
tile_sizeINT51264–4096Tile edge in source pixels: 256 and 512 suit most cards. If the card runs out, the tile is halved and the frame retried.
overlapINT320–1024How far neighbouring tiles overlap, in source pixels: 0 = hard joins; 32 to 64 hides them on most models.
crfFLOAT16.00–51Quality of the written clip, lower is better and larger: 0 = lossless; 16 = default; 23 = typical; 28 = small.
motion_modeloptWAS_MOTION_MODELA flow network from Video Motion Model Loader, which the motion is measured with. Left empty, the built-in texture flow measures it.
ema_vfi_modeloptEMA_VFI_MODELThe interpolation network from EMA-VFI Video Model Loader, which draws frames between two when frame_rate changes. Left empty, the nearer frame repeats. Any rate but double needs an 'ours_t' checkpoint.
precisionoptCOMBOautoWhat the upscale model runs in: 'auto' = half precision where the model declares it safe; '32 bit float' = every model's safest.
new_framesoptCOMBOautoHow a frame between two source frames is made when frame_rate changes: 'auto' = EMA-VFI where the motion accounts for the change, the nearer frame where it does not, such as a mouth changing shape; 'interpolate' = EMA-VFI always; 'blend' = the two mixed; 'hold' = the nearer frame.

Outputs (2)

NameTypeDescription
videoVIDEOThe upscaled clip as an h264 file, with the source's audio. Save Video on 'auto' keeps it without encoding it again.
framesINTHow many frames the upscaled clip holds.