ComfyUI Node

Load Video (C2C)

Frames, audio, mask and VAE latents from one loader

By Code2CollapseΒ·Created 8 months agoΒ·Updated a day agoΒ· 58
Load Video (C2C)
  • vae
  • IMAGE
  • frame_count
  • audio
  • video_info
  • video
  • mask
β—„videoβ–Ύβ–Ί
β—„force_rate0β–Ί
β—„custom_width0β–Ί
β—„custom_height0β–Ί
β—„frame_load_cap0β–Ί
β—„skip_first_frames0β–Ί
β—„select_every_nth1β–Ί
β—„formatNoneβ–Ί

Most video loaders give you frames and call it a day. This one gives you six outputs from a file in ComfyUI's input folder: IMAGE (or LATENT, if you wire a VAE in), frame_count, audio, video_info in VideoHelperSuite's format, a lazy C2C_VIDEO handle, and mask.

That combination is the reason to pick it over the stock loader. Taking audio out of the same node that gives you frames means a restore/upscale pass can hand both to the combiner without a second loader and a second decode. And the C2C_VIDEO handle is the pack's lazy type: it carries the plan - path, in/out points, resampling - and only the consumers that need pixels pay for decoding.

How it works

The file is probed rather than decoded: PyAV reads the container metadata, builds an index table of presentation timestamps, and reads the display matrix so a phone video shot in portrait isn't quietly sideways. Decoding then happens in chunks, respecting your skip/nth/cap selection, and the whole batch is budget-checked against free system memory before it's built - the node refuses with the actual numbers ("this selection is N frames of WxH, X GB as an IMAGE, and only Y GB of RAM is free") and suggests a cap that fits, instead of letting the OS find out for you.

force_rate resamples the output frame rate; 0 keeps the source. custom_width / custom_height resize, and 0 on one of them derives it from the other to preserve aspect. frame_load_cap, skip_first_frames and select_every_nth are the trim trio.

The two non-obvious widgets: format applies a model preset - AnimateDiff, Mochi, LTXV, Hunyuan, Cosmos, Wan - which rounds the frame count onto that model's grid and snaps the dimensions to the model's required multiple. This is genuinely useful, and it's also trimming your clip: pick Wan and a 33-frame selection becomes 33 only if the maths works, otherwise it's capped down to the nearest valid count. If your clip got shorter and you didn't ask, this is the cause. vae is the switch that changes the IMAGE output into LATENT by encoding it on the way out, saving you a separate encode node.

Install

cd ComfyUI/custom_nodes
git clone https://github.com/Code2Collapse/ComfyUI-CustomNodePacks.git

Or "CustomNodePacks" in ComfyUI Manager. This node needs PyAV (av), which ships with modern ComfyUI; if the loader complains, pip install av in ComfyUI's own Python environment. There's no model to download for the loader itself, and the pack's heavy dependencies (SAM, ViTMatte, transformers) belong to other nodes - the loader works fine without them. Don't run the pack's requirements.txt on a working install; it wants to reinstall things ComfyUI already provides, and the README warns about exactly that.

The file list comes from ComfyUI/input/ - it's a widget, not a path field, so a video sitting on your D: drive needs copying there or the sibling Load Video Path (C2C) node.

Where people get burned

The frame grid is the big one. Hand a video model a count it doesn't accept and it doesn't error - it reinterprets the clip and your motion comes out at the wrong speed. The format preset is there to prevent that. If you'd rather keep your exact frame count and pad around it, the pack's AV Handles node does the pad-and-trim properly.

Second: the resize. custom_width and custom_height are the loader's resample, done before any model sees it, and a format preset may change the size again. Two nodes quietly disagreeing about dimensions is a classic source of "why is my output squashed".

Third, and it bites in VRAM terms rather than errors: a big selection of 1080p frames as an IMAGE batch lives in system RAM, not VRAM. The budget check is real, and on a 32 GB machine a long 4K selection will trip it. Lower the cap, raise nth, or resize on the way in.

Category🐺 C2C/🧰 Core/Video

Inputs (9)

NameTypeDefaultDescription
videoCOMBOVideo file from the ComfyUI input folder.
force_rateFLOAT00–240Output FPS; 0 keeps source.
custom_widthINT00–16384Output width; 0 keeps source or derives from height.
custom_heightINT00–16384Output height; 0 keeps source or derives from width.
frame_load_capINT00–1000000Max frames; 0 = no cap.
skip_first_framesINT00–1000000Skip this many frames first.
select_every_nthINT11–100000Keep every Nth frame.
vaeoptVAEEncode to LATENT instead of IMAGE.
formatoptCOMBONoneModel preset (frame grid / dim multiple).

Outputs (6)

NameTypeDescription
IMAGEIMAGE,LATENTβ€”
frame_countINTβ€”
audioAUDIOβ€”
video_infoVHS_VIDEOINFOβ€”
videoC2C_VIDEOβ€”
maskMASKβ€”