Nodes/DIGIT Nodes/DIGIT ElevenLabs Dialogue
ComfyUI Node

DIGIT ElevenLabs Dialogue

Two voices, one node — multi-speaker dialogue straight from a script

By thedepartmentofexternalservices·Created 7 months ago·Updated 2 months ago· 0
DIGIT ElevenLabs Dialogue
    • audio
    ◄text1►
    ◄voice_id1►
    ◄num_entries2►
    ◄modeleleven_v3►
    ◄stability0.50►
    ◄seed1►
    ◄api_key►
    ◄text2►
    ◄voice_id2►
    ◄text3►
    ◄voice_id3►
    ◄text4►
    ◄voice_id4►
    ◄text5►
    ◄voice_id5►
    ◄text6►
    ◄voice_id6►
    ◄text7►
    ◄voice_id7►
    ◄text8►
    ◄voice_id8►
    ◄text9►
    ◄voice_id9►
    ◄text10►
    ◄voice_id10►
    ◄language_code►
    ◄apply_text_normalizationauto►
    ◄output_formatpcm_44100►

    Here's the problem the DIGIT ElevenLabs Dialogue node solves: you need two characters talking to each other, and the naive approach is two text-to-speech nodes, two separate runs, and a pile of manual alignment. This node does it in one shot - you give it up to ten dialogue lines, each assigned to a voice ID, and it returns a single combined audio track. No stitching, no timing math, no "did the pause land right" anxiety.

    It's the multi-speaker member of the pack's ElevenLabs family, and the payoff is most obvious in the obvious place: a scripted conversation - a product ad, a character scene, a two-person narration. The voices are ElevenLabs voices (the commercial bar for this kind of work), so the quality ceiling is high, and the node handles the layout so you don't have to.

    How it works

    The inputs mirror a script: text1 with its voice_id1, and so on up to text10/voice_id10. num_entries (default 2, max 10) tells the node how many lines to actually use, so you can leave the rest of the text fields empty and just flip the count. Each line goes to ElevenLabs with its assigned voice, and the segments come back assembled into one audio output (a ComfyUI AUDIO tensor, PCM 44.1kHz by default).

    The shared voice controls: stability (0.5 default - lower for more expressiveness, higher for consistency), seed for reproducible takes, and model (eleven_v3). There's also language_code for non-English dialogue and apply_text_normalization (auto/on/off) for how numbers and abbreviations are expanded. output_format lets you switch between pcm_44100, mp3_44100_192, and opus_48000_192 - the MP3/Opus options are worth it if you're sending the audio straight to a video encoder.

    The api_key field is optional because the node auto-detects ELEVENLABS_API_KEY (or the pack's DIGIT_ELEVENLABS_API_KEY) from the environment - paste it on the node only if you're not using env vars.

    Installing it

    Standard pack install:

    cd ComfyUI/custom_nodes
    git clone https://github.com/thedepartmentofexternalservices/comfyui-digit.git
    cd comfyui-digit
    pip install -r requirements.txt
    

    Or ComfyUI Manager → search comfyui-digit → install → restart. Then set your key before starting ComfyUI:

    export ELEVENLABS_API_KEY=your_key_here
    

    Where to get the voice IDs

    The natural pairing is the pack's DIGIT ElevenLabs Voice Selector node, which outputs a voice_id you can wire straight into voice_id1, voice_id2, etc. - or paste IDs you already have from the ElevenLabs site. That's the whole trick to this node: the per-line voice assignment is what makes a two-voice track sound like two different people instead of one person doing voices. Set the lines, assign the voices, hit run, and the single audio output is ready for your video or your save node.

    CategoryDIGIT/ElevenLabs

    Inputs (28)

    NameTypeDefaultDescription
    text1STRINGDialogue line 1.
    voice_id1STRINGVoice ID for line 1.
    num_entriesINT21–10Number of dialogue entries to use.
    modelCOMBOeleven_v31 options: eleven_v3
    stabilityFLOAT0.500–1—
    seedINT10–4294967295—
    api_keyoptSTRINGElevenLabs API key. Auto-detected from ELEVENLABS_API_KEY env var.
    text2optSTRING—
    voice_id2optSTRING—
    text3optSTRING—
    voice_id3optSTRING—
    text4optSTRING—
    voice_id4optSTRING—
    text5optSTRING—
    voice_id5optSTRING—
    text6optSTRING—
    voice_id6optSTRING—
    text7optSTRING—
    voice_id7optSTRING—
    text8optSTRING—
    voice_id8optSTRING—
    text9optSTRING—
    voice_id9optSTRING—
    text10optSTRING—
    voice_id10optSTRING—
    language_codeoptSTRING—
    apply_text_normalizationoptCOMBOauto3 options: auto, on, off
    output_formatoptCOMBOpcm_441003 options: pcm_44100, mp3_44100_192, opus_48000_192

    Outputs (1)

    NameTypeDescription
    audioAUDIO—