Load Video (C2C)
Frames, audio, mask and VAE latents from one loader
- vae
- IMAGE
- frame_count
- audio
- video_info
- video
- mask
Most video loaders give you frames and call it a day. This one gives you six outputs from a file in ComfyUI's input folder: IMAGE (or LATENT, if you wire a VAE in), frame_count, audio, video_info in VideoHelperSuite's format, a lazy C2C_VIDEO handle, and mask.
That combination is the reason to pick it over the stock loader. Taking audio out of the same node that gives you frames means a restore/upscale pass can hand both to the combiner without a second loader and a second decode. And the C2C_VIDEO handle is the pack's lazy type: it carries the plan - path, in/out points, resampling - and only the consumers that need pixels pay for decoding.
How it works
The file is probed rather than decoded: PyAV reads the container metadata, builds an index table of presentation timestamps, and reads the display matrix so a phone video shot in portrait isn't quietly sideways. Decoding then happens in chunks, respecting your skip/nth/cap selection, and the whole batch is budget-checked against free system memory before it's built - the node refuses with the actual numbers ("this selection is N frames of WxH, X GB as an IMAGE, and only Y GB of RAM is free") and suggests a cap that fits, instead of letting the OS find out for you.
force_rate resamples the output frame rate; 0 keeps the source. custom_width / custom_height resize, and 0 on one of them derives it from the other to preserve aspect. frame_load_cap, skip_first_frames and select_every_nth are the trim trio.
The two non-obvious widgets: format applies a model preset - AnimateDiff, Mochi, LTXV, Hunyuan, Cosmos, Wan - which rounds the frame count onto that model's grid and snaps the dimensions to the model's required multiple. This is genuinely useful, and it's also trimming your clip: pick Wan and a 33-frame selection becomes 33 only if the maths works, otherwise it's capped down to the nearest valid count. If your clip got shorter and you didn't ask, this is the cause. vae is the switch that changes the IMAGE output into LATENT by encoding it on the way out, saving you a separate encode node.
Install
cd ComfyUI/custom_nodes
git clone https://github.com/Code2Collapse/ComfyUI-CustomNodePacks.git
Or "CustomNodePacks" in ComfyUI Manager. This node needs PyAV (av), which ships with modern ComfyUI; if the loader complains, pip install av in ComfyUI's own Python environment. There's no model to download for the loader itself, and the pack's heavy dependencies (SAM, ViTMatte, transformers) belong to other nodes - the loader works fine without them. Don't run the pack's requirements.txt on a working install; it wants to reinstall things ComfyUI already provides, and the README warns about exactly that.
The file list comes from ComfyUI/input/ - it's a widget, not a path field, so a video sitting on your D: drive needs copying there or the sibling Load Video Path (C2C) node.
Where people get burned
The frame grid is the big one. Hand a video model a count it doesn't accept and it doesn't error - it reinterprets the clip and your motion comes out at the wrong speed. The format preset is there to prevent that. If you'd rather keep your exact frame count and pad around it, the pack's AV Handles node does the pad-and-trim properly.
Second: the resize. custom_width and custom_height are the loader's resample, done before any model sees it, and a format preset may change the size again. Two nodes quietly disagreeing about dimensions is a classic source of "why is my output squashed".
Third, and it bites in VRAM terms rather than errors: a big selection of 1080p frames as an IMAGE batch lives in system RAM, not VRAM. The budget check is real, and on a 32 GB machine a long 4K selection will trip it. Lower the cap, raise nth, or resize on the way in.
Inputs (9)
| Name | Type | Default | Description |
|---|---|---|---|
| video | COMBO | Video file from the ComfyUI input folder. | |
| force_rate | FLOAT | 00β240 | Output FPS; 0 keeps source. |
| custom_width | INT | 00β16384 | Output width; 0 keeps source or derives from height. |
| custom_height | INT | 00β16384 | Output height; 0 keeps source or derives from width. |
| frame_load_cap | INT | 00β1000000 | Max frames; 0 = no cap. |
| skip_first_frames | INT | 00β1000000 | Skip this many frames first. |
| select_every_nth | INT | 11β100000 | Keep every Nth frame. |
| vaeopt | VAE | Encode to LATENT instead of IMAGE. | |
| formatopt | COMBO | None | Model preset (frame grid / dim multiple). |
Outputs (6)
| Name | Type | Description |
|---|---|---|
| IMAGE | IMAGE,LATENT | β |
| frame_count | INT | β |
| audio | AUDIO | β |
| video_info | VHS_VIDEOINFO | β |
| video | C2C_VIDEO | β |
| mask | MASK | β |