IterateThruSpeakers
Who's talking now? Pick a speaker out of any audio, one at a time
- audio
- total_segments
- start_time
- duration
- speaker
The name undersells it. IterateThruSpeakers doesn't isolate anybody - it answers the question "who speaks, and when" for a whole audio file, then hands you the timestamps for one speaker turn at a time. Think of it as the diarization wheel you turn inside a loop: it exists so a single audio file can drive a batch of per-speaker jobs. The author's own demo is old-sitcom audio chopped into segments, each segment turned into a video clip, then all the clips stitched back into one long video.
Audio is still the thinnest layer in ComfyUI - the ecosystem spent its whole life image-first and bolted sound on later, which is exactly why a node like this lives in a small community pack with a heavy dependency stack instead of in core. It's niche, but if your workflow starts from a podcast, a recording, or any dialogue track, it's the piece that turns "audio in" into "a per-speaker pipeline."
What it actually does
The node runs real speaker diarization via pyannote.audio - specifically the community variant of the gated pyannote/speaker-diarization-community-1 pipeline. Under the hood it:
- flattens the incoming
AUDIOto mono and resamples it to 16 kHz (pyannote's preferred rate), - runs the diarization model, which produces a list of
(start, end, speaker)turns, - drops turns under one second, merges back-to-back turns from the same speaker,
- and returns the segment you asked for by
index.
That merge-and-filter step matters. Raw diarization output is noisy - a speaker's line gets split into a dozen fragments, and the model hallucinates half-second blips. By the time this node is done you get clean blocks: SPEAKER_00 talks from 12.3s to 18.0s, SPEAKER_01 from 19.2s to 24.7s, and so on, numbered in the order they speak.
One quirk worth knowing: the duration it returns isn't strictly the speech itself. It extends forward to just before the next speaker starts (minus a 0.1s buffer), so you get the trailing pause included. For feeding video generation that's a feature - nobody wants a clip that cuts off the instant someone stops talking. If you wanted razor-tight clips instead, subtract a little yourself.
The inputs and outputs that matter
Only three inputs, and two are boring:
audio(AUDIO) - whatever you loaded or generated.hf_token(STRING) - your Hugging Face token. This is the gate; without it nothing runs.index(INT, default 1, range 1–100) - which speaker block to return, 1-based. This is the one you'll wire to a loop counter.
Outputs: total_segments (INT, how many blocks the file got split into), start_time (FLOAT), duration (FLOAT), and speaker (STRING, like SPEAKER_00). Note what's not here: no audio output. This node gives you numbers, not clips. Feed start_time and duration into an audio-slicing node (VHS audio tools or any ComfyUI audio trimmer) to actually cut the segment, and use total_segments as the stop condition for your loop. Out-of-range index returns zeros, so read total_segments first and cap the loop there.
Installing it is the fiddly part
The pack's requirements.txt is deliberately gutted - pyannote's dependency tree is huge and the author chose not to auto-install it into your ComfyUI environment, which is also why installing via ComfyUI Manager can look like a failure. The real path:
cd ComfyUI/custom_nodes
git clone https://github.com/trashkollector/ComfyUI-Speaker-Isolation-Community
cd ComfyUI-Speaker-Isolation-Community
pip install "pyannote-audio>=4.0.4"
The README's cautious route is pip install "pyannote-audio>=4.0.4" --no-deps first, then add dependencies as errors surface. You'll also want ffmpeg on your PATH - pyannote and torchaudio lean on it for loading audio.
Then the part everyone trips on: the models are gated. You need a Hugging Face account, a token with read access, and you must accept the user agreements on multiple pages - pyannote/segmentation-3.0, the community diarization pipeline, and the speechbrain embedding model. Accepting one but not the others gets you a download error. First run pulls the models into ~/.cache/huggingface; after that, runs can work offline.
Where people get burned
- 401/403 on load - token wrong, or you haven't accepted terms on all the model pages. The README calls this out hard for a reason.
- The heavy install breaks your env - pyannote pulls torch, torchaudio, lightning, and friends. If your ComfyUI env is precious, consider a dedicated venv before you pip-install into it.
- Recursion errors - pyannote/speechbrain sometimes hit Python's recursion limit; the pack's other node works around it, this one less so. Updating speechbrain usually helps.
It's a small, single-purpose node with a genuinely awkward install. But if you're building a per-speaker video pipeline, there isn't a cleaner native way to get "who talks when" out of an audio file inside ComfyUI.
Inputs (3)
| Name | Type | Default | Description |
|---|---|---|---|
| audio | AUDIO | — | |
| hf_token | STRING | — | |
| index | INT | 11–100 | — |
Outputs (4)
| Name | Type | Description |
|---|---|---|
| total_segments | INT | — |
| start_time | FLOAT | — |
| duration | FLOAT | — |
| speaker | STRING | — |