ComfyUI Node
BFS Shot Planner
Split a long video into model-sized shots (at camera cuts, fixed length, or by hand), give every shot its own reference image and prompt, and run the rest of the graph once per shot. Join the decoded shots with BFS Shot Join.
BFS Shot Planner
- ref_image
- ref_image_2
- vlm
- shots
- count
- fps
- width
- height
- audio
- summary
- timeline
- ref_image
- ref_image_2
◄plan{"video": "", "fps": 24.0, "grid": "H3 (17n+5)", "mode": "shots", "max_s": 4.5, "min_s": 1.0, "sensitivity": 0.5, "max_parts": 0, "max_total_s": 0.0, "bounds": [], "segs": [], "global_ref": "", "global_ref2": "", "global_prompt": "", "megapixels": 0.15, "multiple": 32, "detector": "adaptive", "run": "auto", "filters": {}, "skip_fill": "original", "cast": {}, "cast_assign": true, "cast_split": false, "cast_only": false, "mask_cfg": {}, "vlm_cfg": {}, "audio_mode": "auto", "ref_details": {}}►
◄prompt—►
CategoryBFS/shot loop
Inputs (5)
| Name | Type | Default | Description |
|---|---|---|---|
| plan | STRING | {"video": "", "fps": 24.0, "grid": "H3 (17n+5)", "mode": "shots", "max_s": 4.5, "min_s": 1.0, "sensitivity": 0.5, "max_parts": 0, "max_total_s": 0.0, "bounds": [], "segs": [], "global_ref": "", "global_ref2": "", "global_prompt": "", "megapixels": 0.15, "multiple": 32, "detector": "adaptive", "run": "auto", "filters": {}, "skip_fill": "original", "cast": {}, "cast_assign": true, "cast_split": false, "cast_only": false, "mask_cfg": {}, "vlm_cfg": {}, "audio_mode": "auto", "ref_details": {}} | The panel writes this. JSON with the video, split settings, boundaries and the per-shot reference and prompt. |
| ref_imageopt | IMAGE | Default reference for every shot that has none of its own (overrides the panel's global reference). | |
| ref_image_2opt | IMAGE | Default second reference (e.g. a full-body photo). | |
| promptopt | STRING | Default prompt for shots without their own (overrides the panel's global prompt). | |
| vlmopt | CLIP | Optional vision-language model (CLIPLoader with qwen3vl_4b / qwen3vl_8b). It looks at every shot and suggests what to segment, a description of the shot (fills {shot} in the prompt) and whether to run it. Settings in the panel's VLM card. |
Outputs (10)
| Name | Type | Description |
|---|---|---|
| shots | BFS_SHOT | One item per shot. Every node that receives this list runs once per shot; connect it to BFS Shot Unpack or BFS Shot H3 Conditioning, sample, decode, then BFS Shot Join. |
| count | INT | Number of shots that will run. |
| fps | FLOAT | Timeline frame rate. |
| width | INT | Generation width. |
| height | INT | Generation height. |
| audio | AUDIO | The whole soundtrack, trimmed to the planned duration. |
| summary | STRING | Human-readable plan. |
| timeline | BFS_SHOT_TIMELINE | Every shot in order, including the ones that do not run (disabled or filtered out). Connect it to BFS Shot Join so skipped shots are filled with the original video (or dropped). |
| ref_image | IMAGE | The references the shots use, without repeats: one image when every shot shares the same reference. |
| ref_image_2 | IMAGE | The second references the shots use, without repeats. |