ComfyUI Node

FlashVSR

Fast video upscaling, if you can win the model-folder fight

By wenchengxiang·Created 2 months ago·Updated 9 days ago· 3
FlashVSR
  • frames
  • image
◄modelFlashVSR-v1.1►
◄modetiny►
◄scale2►
◄color_fixtrue►
◄tiled_vaetrue►
◄tiled_dittrue►
◄unload_ditfalse►
◄tile_size256►
◄tile_overlap24►
◄seed0►

What you're actually installing

ComfyUI-Practical-Tools is a grab-bag of routing, loop, mask and image-batch utilities, and FlashVSR is one of the few heavyweight model nodes hiding in it - the pack's README never mentions it at all. What ships is a port of lihaoyun6's ComfyUI-FlashVSR_Ultra_Fast, squashed into one node class, with the auto-downloader deliberately ripped out. That removal is the single most important fact here.

Why you'd reach for it

Upscaling video is not upscaling images N times. Run a per-frame upscaler on a clip and fine repeating texture - wallpaper, brick, fabric - shimmers, because nothing forces frame 4 and frame 5 to resolve the same pattern the same way. That's the defining failure of the job, and why you want a model that sees frames together.

FlashVSR landed in October 2025 and got adopted on speed, not quality: it's a distilled one-step model, so a pass costs about what one sampling step costs. The honest split: FlashVSR when the source is already decent, SeedVR2 when it's genuinely bad (256px garbage climbing to 1024px). FlashVSR's chatter is fading and SeedVR2's isn't, but for "my 540p clip could be nicer" this is the cheap non-rented-GPU answer.

How it works

You feed it an IMAGE batch of sequential frames. It pads the frame count up to a multiple of 8 plus 5 by repeating the last frame - a streaming-model shape requirement - and drops the extras before returning. Then it bicubic-upscales by your scale and center-crops to dimensions divisible by 128, because the DiT works at the scaled size, not your input size.

Sampling is one step at cfg 1.0 - there's no step count to tune. tiny and tiny-long decode latents with a bundled TCDecoder rather than the Wan VAE; full uses the real Wan 2.1 VAE with its encoder stripped, the slow, hungry path.

tiled_dit slices each frame into overlapping tiles, runs the model per tile, and blends them into a full canvas through a feathered mask so seams don't show.

The inputs worth touching

  • mode - tiny (default), tiny-long, full. The author's tooltip says tiny-long "significantly reduce[s] VRAM used with long video input"; start there for longer clips.
  • scale - integer, default 2. The bicubic pre-upscale means the model does less work than you'd think at high values.
  • tiled_dit, tile_size (default 256), tile_overlap - your VRAM knobs. Wrinkle: tile_size's tooltip claims 0 disables tiling, but the widget floor is 32; turn it off with the tiled_dit boolean instead.
  • color_fix - on by default; leave it on, generative upscalers drift colour on video.

unload_dit frees the DiT before decode to shave the VRAM peak, at a speed cost; tiled_vae is the decode-side equivalent, where disabling it is faster and hungrier. Output is one IMAGE batch named image, the same frame count you fed in.

Installing it

Manager → search Practical-Tools, or:

cd ComfyUI/custom_nodes
git clone https://github.com/wenchengxiang/ComfyUI-Practical-Tools.git

Restart ComfyUI. The pack's requirements.txt (onnxruntime, nvidia-vfx, openai, gguf) belongs to its other nodes - this node needs nothing your install lacks.

Then the models, which nothing fetches for you. Make the folder name exactly what the model dropdown says:

ComfyUI/models/FlashVSR-v1.1/
├── diffusion_pytorch_model_streaming_dmd.safetensors
├── Wan2.1_VAE.pth
├── LQ_proj_in.ckpt
└── TCDecoder.ckpt

The default is FlashVSR-v1.1, so that's the folder name you need; switching to FlashVSR sends it to models/FlashVSR/ instead. Weights come from JunhaoZhuang's FlashVSR and FlashVSR-v1.1 repos on Hugging Face; the prompt embedding (posi_prompt.pth) ships inside the pack, and all four files are existence-checked even in tiny mode.

When it goes wrong

"Model directory does not exist: .../models/FlashVSR-v1.1" is the number one hit, and it's almost always the naming thing above - files in models/FlashVSR/, dropdown on v1.1. Rename the folder or change the dropdown.

OOM. tiled_dit on plus tile_size down to 128–192 is the lever. Tiling isn't free: the best-known account from FlashVSR's original thread runs 81 frames of 640x880 at 2x on a 24GB 3090 with both DiT and VAE tiling off - fast, filling most of the card - while tiling took that under a third of 24GB. On 8–12GB, expect a coffee.

"No devices found to run FlashVSR!" - the node only accepts CUDA or Apple MPS. CPU-only installs are out.

Slow repeat runs. The node rebuilds its pipeline every execution instead of caching it, so each queue re-reads the weights from disk.

One-frame input does something odd. A single image comes back as the temporal median across the padded frames - a side effect of the padding, not a supported still-image mode.

CategoryPractical-Tools/SResolution

Inputs (11)

NameTypeDefaultDescription
framesIMAGESequential video frames as IMAGE tensor batch
modelCOMBOFlashVSR-v1.1Model version.
modeCOMBOtinyUsing "tiny-long" mode can significantly reduce VRAM used with long video input.
scaleINT21–9223372036854776000—
color_fixBOOLEANtrueUse color fix algorithm to preserve original colors.
tiled_vaeBOOLEANtrueDisable tiling: faster decode but higher VRAM usage.
tiled_ditBOOLEANtrueSignificantly reduces VRAM usage at the cost of speed.
unload_ditBOOLEANfalseUnload DiT before decoding to reduce VRAM peak at the cost of speed.
tile_sizeINT25632–1024DiT tile size; 0 disables tiling.
tile_overlapINT248–512—
seedINT00–1125899906842624—

Outputs (1)

NameTypeDescription
imageIMAGE—