CineSpatial · TranscribeAlignDialogue
Transcribe dialogue and pin every word to a timestamp
- artifact_json
Plain transcription gets you words with a vague "around here." This node gets you a dialogue_alignment artifact - the transcript plus word-level timing - which is the piece of data that makes everything else in the CineSpatial pack possible. DetectActiveSpeaker can't tell you who's talking without knowing when each line lands, and that timing comes from here.
How it works
Like every worker node in the pack, it's a thin client. It takes your dialogue_audio_path, asks the fixed loopback runner at http://127.0.0.1:8199/v1 for the transcribe_align_dialogue operation, and the actual model - the "whisper" slot in the pack's health report - runs in the separate cinespatial runner service, not inside ComfyUI. The runner writes its result into a fresh per-run directory, verifies the files (non-zero length, SHA-256), and returns a dialogue_alignment artifact. The node converts the paths to downloadable worker_ref entries in ComfyUI's output folder, and hands you the whole thing as artifact_json.
The alignment part is the part that isn't just ASR. Word-level timestamps are what let a downstream node correlate speech with the video track, or with an expected script, instead of hoping for the best.
Inputs and outputs
The required input is dialogue_audio_path (STRING) - the file you want transcribed. That's it for the minimum. But there are two optional inputs worth understanding:
expected_language(STRING) - if you know the spoken language, tell it. A hint here steers recognition and avoids the model burning effort guessing between similar languages.expected_text(STRING) - if you already have the script (subtitles, a shooting script, a caption file), pass it in. The node can then align against known text rather than transcribing blind, which is far more reliable for checking that what was said matches what's on paper.
Leave both empty and it just transcribes and aligns from scratch. The output is a single artifact_json (STRING, and this node is an output node so it prints). That string is exactly what CineSpatialDetectActiveSpeaker wants on its dialogue_artifact_json input - you can literally wire the socket across.
Install and the gotcha
Same pack, same story: ComfyUI Manager Git URL installer, or
cd ComfyUI/custom_nodes
git clone https://github.com/Vighneshjs/ComfyUI-CineSpatial
then restart. Nothing to pip-install (requirements.txt is empty), no weights to download here - the transcription model lives in the separate cinespatial runner, and if that isn't running you'll get CineSpatial runner service is unavailable or invalid. The health node tells you whether the whisper backend is actually ready before you bother.
Troubleshooting
- Garbage transcript on a noisy clip → feed it the dialogue stem from
CineSpatialBanditSeparate, not the full mix. Separation first is the intended pipeline for a reason. expected_textmakes things worse → only pass text you're confident is what's actually spoken. If the clip was re-edited, ADR'd, or cut down and your script is stale, alignment against wrong text can produce a confident-sounding mess. When in doubt, transcribe blind and compare.- Long files → operations have a 3600-second timeout, so a feature-length dialogue track won't hard-timeout, but the node blocks while the runner works. There's no progress bar in the graph; plan the run accordingly.
- Schema errors → the runner answered but returned a different
schema_versionthan the pack's0.1. Update the runner package.
Start here, hand the artifact_json to DetectActiveSpeaker, and you've got speaker-tagged dialogue - which is a surprisingly long way down the road to a usable cinematic analysis.
Inputs (3)
| Name | Type | Default | Description |
|---|---|---|---|
| dialogue_audio_path | STRING | — | |
| expected_languageopt | STRING | — | |
| expected_textopt | STRING | — |
Outputs (1)
| Name | Type | Description |
|---|---|---|
| artifact_json | STRING | — |