🎬 智绘_批量视频加载器
Pull frames and audio out of any video in a folder, one index at a time
- images
- audio
- video_info
- filename_text
Video-to-frame extraction is a ComfyUI rite of passage, and ZH_BatchVideoLoader (🎬 智绘_批量视频加载器) is the 智绘灵箱 pack's entry in that category - with a twist that most frame loaders skip: it gives you the audio track too. It's a folder scanner that picks a video by index (video 0, video 1, …), pulls frames out with OpenCV, grabs the audio with PyAV, and hands you an IMAGE batch, an AUDIO tensor, and a VIDEO_INFO structure that keeps your video-aware downstream nodes honest about FPS and frame counts. If you're doing video-to-video, frame interpolation prep, or anything that needs the original timing, this is the loader to wire first.
How it works
The engine split is worth knowing: frames via OpenCV (fast, and OpenCV is bundled with ComfyUI's video stack), audio via PyAV (the "VHS-style" compatibility layer - the source calls it 兼容性王道, compatibility first). PyAV is normally present in modern ComfyUI installs; if it isn't, the node degrades gracefully and just returns no audio rather than crashing - the source literally prints a warning and continues.
The video is picked by video_index from whatever's in folder_path (sorted). Then you shape the extraction:
force_rate- 0 keeps the source FPS; any other value (up to 120) forces the output frame rate. For slowing a clip's frame cadence or matching a target FPS.resize_width/resize_height- 0 = original dimensions, else it resizes the extracted frames (steps of 8).frame_load_cap- 0 = unlimited; set a cap to stop after N frames.skip_first_frames- chop the head off the clip.frame_interval- extract every Nth frame. Interval 2 halves the frame count, which is your cheap "temporal downscale."
The outputs that matter
Four sockets, and they map to real downstream needs: images (IMAGE) for any vision node; audio (AUDIO) for the audio nodes in your graph; video_info (VIDEO_INFO) - a structured bundle holding initial/loaded FPS, frame counts, dimensions and duration, which is what VIDEO_INFO-aware nodes consume to stay in sync; and filename_text (STRING) for saving with the right name. If your video nodes don't take VIDEO_INFO, you can usually ignore it - but it's the reason this node composes cleanly with the pack's other video tools.
Install
Part of the 智绘灵箱 (ComfyUI-ZhiHui) pack:
cd ComfyUI/custom_nodes
git clone https://github.com/zhuyungen/ComfyUI-ZhiHui.git
Restart ComfyUI (or ComfyUI Manager, "智绘灵箱" / "ComfyUI-ZhiHui"). Needs OpenCV (opencv-python, in the pack's requirements) and ideally PyAV - which the pack's README doesn't even mention, so if audio comes back empty, that's the missing piece, not a bug in your graph.
Where people get burned
- Audio requires PyAV, and the README won't warn you. The requirements file lists
opencv-pythonbut notav; the code falls back silently. If you need audio,pip install avexplicitly. frame_intervalandforce_rateare different levers. Interval drops frames (you lose content); force_rate re-times what's there (you lose cadence, keep content). Reaching for the wrong one is the classic mistake.- Folder order defines
video_index. It's alphabetical by filename, so video 0 might not be the one you think. Check thefilename_textoutput before trusting the index. VIDEO_INFOis a pack convention, not a universal type. A workflow full of non-VIDEO_INFO nodes won't break - you just won't use that socket.
For "which video in this folder do I process next" batch work, the index-based selection is genuinely handy - loop the index, keep everything else the same, and every video in the folder gets processed with identical settings. That's the workflow this node was built for.
Inputs (8)
| Name | Type | Default | Description |
|---|---|---|---|
| folder_path | STRING | — | |
| video_index | INT | 00–999999 | — |
| force_rate | INT | 00–120 | 0=原速。设置数值可强制改变输出FPS |
| resize_width | INT | 00–4096 | 0=原宽 |
| resize_height | INT | 00–4096 | 0=原高 |
| frame_load_cap | INT | 00–10000 | 0=无限制 |
| skip_first_frames | INT | 00–10000 | — |
| frame_interval | INT | 11–100 | 间隔取帧 |
Outputs (4)
| Name | Type | Description |
|---|---|---|
| images | IMAGE | — |
| audio | AUDIO | — |
| video_info | VIDEO_INFO | — |
| filename_text | STRING | — |