ComfyUI Node

IterateThruSpeakers

Who's talking now? Pick a speaker out of any audio, one at a time

By trashkollector·Created 6 months ago·Updated 6 months ago· 0
IterateThruSpeakers
  • audio
  • total_segments
  • start_time
  • duration
  • speaker
◄hf_token►
◄index1►

The name undersells it. IterateThruSpeakers doesn't isolate anybody - it answers the question "who speaks, and when" for a whole audio file, then hands you the timestamps for one speaker turn at a time. Think of it as the diarization wheel you turn inside a loop: it exists so a single audio file can drive a batch of per-speaker jobs. The author's own demo is old-sitcom audio chopped into segments, each segment turned into a video clip, then all the clips stitched back into one long video.

Audio is still the thinnest layer in ComfyUI - the ecosystem spent its whole life image-first and bolted sound on later, which is exactly why a node like this lives in a small community pack with a heavy dependency stack instead of in core. It's niche, but if your workflow starts from a podcast, a recording, or any dialogue track, it's the piece that turns "audio in" into "a per-speaker pipeline."

What it actually does

The node runs real speaker diarization via pyannote.audio - specifically the community variant of the gated pyannote/speaker-diarization-community-1 pipeline. Under the hood it:

  • flattens the incoming AUDIO to mono and resamples it to 16 kHz (pyannote's preferred rate),
  • runs the diarization model, which produces a list of (start, end, speaker) turns,
  • drops turns under one second, merges back-to-back turns from the same speaker,
  • and returns the segment you asked for by index.

That merge-and-filter step matters. Raw diarization output is noisy - a speaker's line gets split into a dozen fragments, and the model hallucinates half-second blips. By the time this node is done you get clean blocks: SPEAKER_00 talks from 12.3s to 18.0s, SPEAKER_01 from 19.2s to 24.7s, and so on, numbered in the order they speak.

One quirk worth knowing: the duration it returns isn't strictly the speech itself. It extends forward to just before the next speaker starts (minus a 0.1s buffer), so you get the trailing pause included. For feeding video generation that's a feature - nobody wants a clip that cuts off the instant someone stops talking. If you wanted razor-tight clips instead, subtract a little yourself.

The inputs and outputs that matter

Only three inputs, and two are boring:

  • audio (AUDIO) - whatever you loaded or generated.
  • hf_token (STRING) - your Hugging Face token. This is the gate; without it nothing runs.
  • index (INT, default 1, range 1–100) - which speaker block to return, 1-based. This is the one you'll wire to a loop counter.

Outputs: total_segments (INT, how many blocks the file got split into), start_time (FLOAT), duration (FLOAT), and speaker (STRING, like SPEAKER_00). Note what's not here: no audio output. This node gives you numbers, not clips. Feed start_time and duration into an audio-slicing node (VHS audio tools or any ComfyUI audio trimmer) to actually cut the segment, and use total_segments as the stop condition for your loop. Out-of-range index returns zeros, so read total_segments first and cap the loop there.

Installing it is the fiddly part

The pack's requirements.txt is deliberately gutted - pyannote's dependency tree is huge and the author chose not to auto-install it into your ComfyUI environment, which is also why installing via ComfyUI Manager can look like a failure. The real path:

cd ComfyUI/custom_nodes
git clone https://github.com/trashkollector/ComfyUI-Speaker-Isolation-Community
cd ComfyUI-Speaker-Isolation-Community
pip install "pyannote-audio>=4.0.4"

The README's cautious route is pip install "pyannote-audio>=4.0.4" --no-deps first, then add dependencies as errors surface. You'll also want ffmpeg on your PATH - pyannote and torchaudio lean on it for loading audio.

Then the part everyone trips on: the models are gated. You need a Hugging Face account, a token with read access, and you must accept the user agreements on multiple pages - pyannote/segmentation-3.0, the community diarization pipeline, and the speechbrain embedding model. Accepting one but not the others gets you a download error. First run pulls the models into ~/.cache/huggingface; after that, runs can work offline.

Where people get burned

  • 401/403 on load - token wrong, or you haven't accepted terms on all the model pages. The README calls this out hard for a reason.
  • The heavy install breaks your env - pyannote pulls torch, torchaudio, lightning, and friends. If your ComfyUI env is precious, consider a dedicated venv before you pip-install into it.
  • Recursion errors - pyannote/speechbrain sometimes hit Python's recursion limit; the pack's other node works around it, this one less so. Updating speechbrain usually helps.

It's a small, single-purpose node with a genuinely awkward install. But if you're building a per-speaker video pipeline, there isn't a cleaner native way to get "who talks when" out of an audio file inside ComfyUI.

Categoryaudio

Inputs (3)

NameTypeDefaultDescription
audioAUDIO—
hf_tokenSTRING—
indexINT11–100—

Outputs (4)

NameTypeDescription
total_segmentsINT—
start_timeFLOAT—
durationFLOAT—
speakerSTRING—