Nodes/flashvsr-sm89-ops/FlashVSR Video Upscale (sm89, local)
ComfyUI Node

FlashVSR Video Upscale (sm89, local)

ComfyUI's Cloud FlashVSR Node, Without the Credits or the Upload

By aireet·Created 7 days ago·Updated about 22 hours ago· 3
FlashVSR Video Upscale (sm89, local)
  • video
  • VIDEO
◄target_resolution▾►

The free version of a node you were paying per call for

ComfyUI ships a FlashVSR node that doesn't run FlashVSR - it calls WaveSpeed, which means an account, prepaid credits, and your clip leaving the machine. This pack is the same node running locally: VIDEO in, VIDEO out, between the core Load Video and Save Video nodes, no key and no per-call billing.

Worth knowing where FlashVSR sits: video is the third upscaling job - more pixels over time, where every frame has to agree with its neighbours - and the least settled. FlashVSR broke through in October 2025 on speed, not restoration quality, and the community's mid-2026 line is FlashVSR for a decent source you want done now, SeedVR2 for footage that's genuinely low-resolution. Feed it a mush of a VHS rip and nothing here will save you.

How it works

Underneath is FlashVSR's Wan2.1-1.3B DiT with LCSA sparse attention and a TCDecoder, run as a one-step DMD restorer over a window of frames so the output stays temporally consistent - the reason you don't just run an image upscaler 24 times a second, where fine repeating texture shimmers because consecutive frames resolve it differently.

Your frames are bicubic-upscaled ×4, center-cropped to a 128px grid, and streamed through in blocks of 8, so the frame count gets trimmed to the 8n+1 rule. The interesting engineering is memory: upstream renders the whole canvas in one pass, so VRAM scales with output pixels and 4K asks for ~45 GiB - that's the OOM. This node keeps anything up to ~2.2 MP untiled, and above that renders overlapping spatial tiles with 256px blend ramps, uploading one tile's crop at a time while the low-res canvas stays in CPU RAM. Cost: about +12% wall time at 2K.

Then there's the operator pack the repo is named for: FP8 linears, fused RMSNorm+RoPE and AdaLN kernels, a channels_last decoder, and a Triton stand-in for mit-han-lab's block-sparse-attention CUDA extension, so you compile nothing. The author's own A/B: ~1.3× over the stock pipeline on a 4090 D.

The two inputs and the one output

Only two inputs, and one of them you'll set once:

  • video - a VIDEO, from Load Video.
  • target_resolution - combo of 720p / 1080p / 2K / 4K, default 1080p. Roughly how tall the result should be.

Output is a single VIDEO - wire it into Save Video and you're done.

Two things surprise people. The output snaps to the model's 128px grid and rounds down, so a 16:9 "1080p" job lands at 1920×1024 - expected, not a bug. And the VRAM peaks in the tooltip are non-monotonic on purpose, for a ~3 second clip: 720p ≈ 9 GB, 1080p ≈ 19 GB, 2K ≈ 21 GB, 4K ≈ 20 GB, because the big buckets tile and the small ones don't. Tiles auto-size to your free VRAM too, so a 16 GB card tiles 1080p instead of dying.

Everything else is fixed on purpose: 1 step, CFG 1.0, seed 0, and a per-tile topk_ratio rescaled so a small tile runs at the same operating point as a full-frame one. Nothing to tune.

Audio survives only when the output frame count matches the input count. With the padding and the 8n+1 trim, that often isn't true, so padded clips come back silent.

Installing it

cd ComfyUI/custom_nodes
git clone https://github.com/aireet/flashvsr-sm89-ops.git

Restart ComfyUI. There's no pip step - the node imports the operator pack from its own checkout. In Manager it's Install from Git with that same URL, or search the pack title.

The first run provisions itself: it clones OpenImagingLab/FlashVSR into the pack folder and pulls ~6.5 GB of v1.1 weights into ComfyUI/models/flashvsr/. Don't kill it halfway - and if git isn't installed the clone fails with instructions instead of guessing.

Runtime needs torch ≥ 2.6 and Triton ≥ 3.2, plus a C compiler and your interpreter's dev headers (sudo apt install python3-dev) because Triton JIT-compiles on first use.

When it goes wrong

fatal error: Python.h: No such file or directory on the first run is that missing dev-header, not a broken install. It compiles once and caches.

The first run is slow. Triton compiles inside it - 9.5 FPS cold versus 13.7 warm on the same clip, by the author's measurement. Judge nothing until run two.

Out of memory. The node raises with the measured numbers and points you at Trim Video; VRAM scales with frame count as much as resolution.

You're not on Ada. The FP8 path wants sm_89 tensor cores (4090/4090 D, L40, RTX 6000 Ada); the bf16 Triton kernels and LCSA run on anything sm_80+, so a 3090 is fine. If the FP8 conversion misbehaves on your architecture, set FS89_FP8= (empty) to skip it and keep the fused norms and channels_last.

Last thing: searching the community corpus for this pack's name returns nothing, so every number in that README is the author's own A/B on a 4090 D - with the raw JSON in benchmarks/, which is more rigor than most packs manage and still one person's card. Worth remembering that FlashVSR's own footprint in the corpus is a tenth of what it was at launch.

Categoryvideo/upscaler

Inputs (2)

NameTypeDefaultDescription
videoVIDEOThe low-resolution clip. Very short clips are padded by holding their last frame; audio is kept when the frame count is unchanged.
target_resolutionCOMBOApproximate output height. The output snaps to the model's 128px grid; larger canvases render as overlapping spatial tiles, so VRAM stays at one tile's cost. Measured peaks for a ~3 s clip: 720p ≈ 9 GB, 1080p ≈ 19 GB, 2K ≈ 21 GB, 4K ≈ 20 GB — all inside a 24 GB card, and tiles auto-size down to fit smaller cards.

Outputs (1)

NameTypeDescription
VIDEOVIDEO—