Nodes/ComfyUI-CineSpatial/CineSpatial · GroundVisualSources
ComfyUI Node

CineSpatial · GroundVisualSources

Type 'the person in the red jacket' and get a grounded result on the video

By Vighneshjs·Created about a month ago·Updated about a month ago· 0
CineSpatial · GroundVisualSources
    • artifact_json
    ◄video_path►
    ◄source_queries_json[]►

    Most of this pack is about the sound. This node is the one that looks at the picture. Give it a video and a list of text queries - "the person in the red jacket", "the detective's desk lamp" - and it grounds those descriptions in the actual frames, producing a visual_evidence artifact that ties your script's words to what's really on screen.

    How it works

    Another thin client to the fixed loopback runner at http://127.0.0.1:8199/v1, this time for the ground_visual_sources operation. The two model slots the health report tracks for it - grounding_dino and sam - are the classic pairing from the image-detailing world: Grounding DINO is open-vocabulary detection (you name it, it returns a bounding box), and SAM turns that box into a precise segmentation mask. Together it's the "Grounded SAM" pattern the KB's masking-detection-detailing doc covers: find the thing you named, then refine the box into an actual region.

    The runner executes in its own isolated environment, writes the evidence to a fresh per-run directory, verifies the files (non-zero length, SHA-256), and the node converts the paths into downloadable worker_ref entries in ComfyUI's output folder. Everything comes back in one artifact_json output.

    The input that matters

    • video_path (STRING, required) - the footage. Absolute path or a filename relative to ComfyUI's input directory.
    • source_queries_json (STRING, required, multiline, defaults to []) - a JSON array of text queries. This is the input you'll actually spend time on. Each query is a thing you're looking for, and the quality of the grounding is almost entirely a function of how specific you make these. "car" will find cars; "the dented red sedan in the driveway" will find the car.

    Output: artifact_json (STRING), an output node so it prints when run.

    Install and the gotcha

    The standard pack install - ComfyUI Manager Git URL installer, or:

    cd ComfyUI/custom_nodes
    git clone https://github.com/Vighneshjs/ComfyUI-CineSpatial
    

    then restart. No pip dependencies (requirements.txt is empty) and no weights download through this pack - the detection and segmentation models live in the separate cinespatial runner. If it isn't running, you get CineSpatial runner service is unavailable or invalid, and CineSpatialWorkerHealth tells you whether grounding_dino and sam are actually loaded.

    Troubleshooting

    • Boxes tighter than you expect - this is the known Grounding DINO behavior: it hugs the object, sometimes cutting off edges. If you're feeding the evidence downstream for cropping or masking, pad the regions before using them. The KB's detailing doc flags the same caveat for grounded detection generally.
    • Vague queries → misses - "person" is a guess; "person in a red jacket" is a detection. Tighten the query before you blame the model.
    • Slow on long videos - detection plus segmentation per frame is the expensive end of this pack. Scope the queries tightly and accept that feature-length footage takes a while; the operation timeout is a generous 3600 seconds, so it'll block the graph rather than fail.
    • Schema errors - the runner answered with a different schema_version than the pack's 0.1; update the runner.

    When it works, you've got text-grounded regions across your footage - which is exactly the bridge between what your script says and what the camera actually shows.

    CategoryCineSpatial/AI Worker

    Inputs (2)

    NameTypeDefaultDescription
    video_pathSTRING—
    source_queries_jsonSTRING[]—

    Outputs (1)

    NameTypeDescription
    artifact_jsonSTRING—