Nodes/ComfyUI-CineSpatial/CineSpatial · DetectActiveSpeaker
ComfyUI Node

CineSpatial · DetectActiveSpeaker

Find out who's actually talking on screen, and when

By Vighneshjs·Created about a month ago·Updated about a month ago· 0
CineSpatial · DetectActiveSpeaker
    • artifact_json
    ◄video_path►
    ◄dialogue_artifact_json—►

    A transcript tells you words and times. It doesn't tell you who said them - and in a two-shot, audio alone can't either. That's the gap this node fills: it takes the aligned dialogue plus the actual video, and produces active_speaker_evidence - which character is on screen and visibly speaking at each moment.

    How it works

    Like the rest of the worker nodes, it's a thin client to the fixed loopback runner at http://127.0.0.1:8199/v1, this time asking for the detect_active_speaker operation. The model - the "talknet" slot in the pack's health report - lives in the separate cinespatial runner service. TalkNet-style active-speaker detection works because it has both things at once: the word timing from the alignment and the video frames, so it can correlate lip motion and on-screen faces with the aligned speech. Audio-only analysis can't do that, which is why this node requires both inputs rather than guessing.

    The runner writes its results into a fresh per-run directory, verifies them (non-zero length, SHA-256), and returns the two required artifact roles for this operation: dialogue_alignment (echoed back) and active_speaker_evidence. Paths come back as downloadable worker_ref entries in ComfyUI's output folder, all wrapped in the node's single artifact_json output.

    The inputs that matter

    • video_path (STRING, required) - the footage. Absolute path or a filename relative to ComfyUI's input directory both work.
    • dialogue_artifact_json (STRING, required, multiline) - this is the artifact_json output of CineSpatialTranscribeAlignDialogue. You can paste it in, or wire the STRING socket directly from one node to the other; the connection is just a JSON string passing between two nodes.

    Two required inputs is actually the whole point of the node. If you hand it a transcript that doesn't match the video's audio - a different take, a re-cut, ADR from a different session - the active-speaker evidence will be confidently wrong. Keep the alignment artifact tied to the exact same source as the video.

    Install and the gotcha

    The usual pack install: ComfyUI Manager Git URL installer, or

    cd ComfyUI/custom_nodes
    git clone https://github.com/Vighneshjs/ComfyUI-CineSpatial
    

    then restart. No pip dependencies (requirements.txt is empty), no weights to download through this pack - the detection model runs in the separate cinespatial runner. If that service isn't up, you get CineSpatial runner service is unavailable or invalid. CineSpatialWorkerHealth will tell you whether the talknet backend is actually ready.

    Troubleshooting

    • "runner service is unavailable or invalid" → the loopback runner isn't running on 8199. Start it (cinespatial-runner --config /opt/cinespatial/runner.json) and check health first.
    • Wrong speaker attributed → almost always a source mismatch: the alignment and the video are from different takes or cuts. Re-run transcription on the actual audio of the video you're feeding here.
    • Schema/operation mismatch errors → the dialogue_artifact_json you fed it isn't a valid 0.1 alignment artifact (maybe it got mangled by a text node in between). Run it through CineSpatialValidateArtifact - that's the local sanity check for exactly this.
    • Feature-length runs block the graph → the operation timeout is a generous 3600 seconds, so it won't hard-fail, but it will hold the workflow while the runner works.

    When this node and the transcript line up, you have speaker-tagged dialogue with timing - the thing downstream pipelines (subtitles per character, who-speaks-when stats, re-synthesis targeting) are all built on.

    CategoryCineSpatial/AI Worker

    Inputs (2)

    NameTypeDefaultDescription
    video_pathSTRING—
    dialogue_artifact_jsonSTRING—

    Outputs (1)

    NameTypeDescription
    artifact_jsonSTRING—