Nodes/ComfyUI-CineSpatial/CineSpatial · ClassifyAudioEvents
ComfyUI Node

CineSpatial · ClassifyAudioEvents

Know exactly what's in the effects track before you build a foley pass

By Vighneshjs·Created about a month ago·Updated about a month ago· 0
CineSpatial · ClassifyAudioEvents
    • artifact_json
    ◄effects_audio_path►
    ◄candidate_intervals_json[]►

    Separating the soundtrack is step one, but a clean effects stem is just audio - you still don't know what's in it. This node answers that. It classifies the sound events in an effects track and returns an effects_evidence artifact: the "door slam at 00:14, glass shatter at 01:02" survey that turns a raw stem into something a sound designer can actually work from.

    How it works

    Same architecture as every worker node here: a thin client to the fixed loopback runner at http://127.0.0.1:8199/v1, asking for the classify_audio_events operation. The audio event classification model - the "beats" slot in the pack's health report, so named after the BEATs audio-pretraining family - runs in the separate cinespatial runner service in its own environment, never inside ComfyUI. Files get written to a fresh per-run directory, verified non-zero-length and SHA-256-checked, and come back as worker_ref references in ComfyUI's output folder. The node's single output is the whole thing as artifact_json.

    Inputs and outputs

    The required input is effects_audio_path (STRING) - the effects stem you want surveyed. If you're following the pack's natural pipeline, that's the effects output from CineSpatialBanditSeparate.

    The one optional input is worth knowing about: candidate_intervals_json (STRING, multiline, defaults to []). If you already know the moments you care about - from a spotting session, your own timeline, or another detector - you can hand it a JSON array of time intervals and the classification scopes itself to those windows instead of chewing through the whole track. Leave it as [] to scan everything.

    Output: artifact_json (STRING), and it's an output node so it prints when run. Downstream, this is the evidence that feeds a re-synthesis pass or a foley-matching step - you now know not just that something happened, but what it was and when.

    Install and the gotcha

    ComfyUI Manager Git URL installer, or:

    cd ComfyUI/custom_nodes
    git clone https://github.com/Vighneshjs/ComfyUI-CineSpatial
    

    then restart. Nothing to pip-install - requirements.txt is a comment - and no model weights download through this pack. The classification model lives in the separate cinespatial runner; if that's not running you'll see CineSpatial runner service is unavailable or invalid. CineSpatialWorkerHealth shows whether the beats backend is loaded before you waste a run.

    Troubleshooting

    • Weak or noisy results → you fed it the full mix instead of a separated effects stem. The classification is much cleaner on the stem; do BanditSeparate first.
    • Intervals that return nothing useful → check your candidate_intervals_json formatting and that the times are inside the actual clip. The pack fails explicitly on missing files, so a silent miss usually means a bad interval or a query that doesn't match anything, not a crash.
    • Long tracks → the operation timeout is a generous 3600 seconds, so it won't hard-fail, but the node blocks the graph while the runner works. Scope with candidate_intervals_json if you only need a few moments.
    • Schema errors → the runner answered with a different schema_version than the pack's 0.1; update the runner to match.

    The move here is: separate, then classify, then decide. You don't guess where the effects are, you read them off the evidence.

    CategoryCineSpatial/AI Worker

    Inputs (2)

    NameTypeDefaultDescription
    effects_audio_pathSTRING—
    candidate_intervals_jsonoptSTRING[]—

    Outputs (1)

    NameTypeDescription
    artifact_jsonSTRING—