Nodes/ComfyUI-KIE-Nodes-Next/Gemini 3.8 Flash Text to speech
ComfyUI Node

Gemini 3.8 Flash Text to speech

Multi-speaker dialogue in one shot

By felipederosilva·Created about a month ago·Updated 6 days ago· 2
Gemini 3.8 Flash Text to speech
    • audio
    • url
    • all_urls_json
    • task_id
    • raw_json
    • credits_consumed
    • credits_left
    ◄speakers[ { "speaker_id": "Speaker 2", "voice_name": "Fenrir", "audio_profile": "ea Lorem sed", "accent": "Transatlantic", "style": "Whisper", "pace": "Natural" }, { "speaker_id": "Speaker 1", "voice_name": "Rasalgethi", "audio_profile": "ut Ut consectetur tempor", "accent": "British (Brixton)", "style": "Whisper", "pace": "Staccato" } ]►
    ◄dialogue_turns[ { "speaker_id": "Speaker 1", "text": "Tero universe ad tremo utroque depulso valde molestiae antea. Certe deficio quidem cupio. Videlicet utrum sursum quas abduco conturbo currus conforto. Ustulo vereor abstergo adfectus arcus carpo subito careo. Voluntarius defluo virtus voluptatibus varius cornu tutamen canis vilis. Auctor adsuesco calcar bellicus dens aggredior aperio testimonium." } ]►
    ◄temperature1.49►
    ◄sceneminim mollit ullamco in►
    ◄sample_contextanim veniam Lorem voluptate exercitation►
    ◄timeout_seconds1200►
    ◄callback_url►

    What it is

    A text-to-speech node that generates conversation, not narration: a cast of speakers defined once, a script of dialogue turns, and one AUDIO object back containing the whole exchange. It calls google/gemini-3-8-flash-tts on KIE, which relays Google's Gemini TTS.

    The use case is specific and worth stating: you have a scene with two or three characters and you want their lines performed with distinct voices. Doing that with a single-voice TTS means stitching files and hand-matching tone. This does it in one job.

    How it works

    The widget shape is the whole node:

    • speakers - a JSON array of speaker configs. Each has speaker_id, voice_name (Gemini voices: Fenrir, Rasalgethi, Iapetus, Schedar…), and then audio_profile, accent, style, pace to describe delivery.
    • dialogue_turns - a JSON array of {speaker_id, text}, performed in order.
    • scene - the room and recording context, e.g. a quiet room with a fireplace crackling.
    • sample_context - the overall tone and delivery reference, e.g. audiobook narration, gentle and inviting.
    • temperature - 0 to 2. Tightens or loosens the performance.

    The join key between the two arrays is speaker_id. Get it wrong and a line comes back in the wrong voice, or nothing comes back at all.

    Under the hood it's an async KIE market task: submitted to /api/v1/jobs/createTask, polled until complete or until timeout_seconds (default 1200 s) elapses. Hence the task_id output, and hence the fact that a job can finish on KIE's side after your node has stopped waiting.

    First thing to do: clear the placeholder defaults. The speakers and dialogue_turns values that ship in the widget are copied from KIE's docs example - two speakers with lorem-ipsum lines and a British (Brixton) accent. The node doesn't care that it's nonsense. It will synthesise it and charge you.

    Inputs and outputs

    Required: speakers, dialogue_turns. Optional: temperature, scene, sample_context, timeout_seconds, callback_url (leave empty in ComfyUI; it's for server-side callbacks).

    Outputs: audio, url, all_urls_json, task_id, raw_json, credits_consumed, credits_left. The native audio output is what you wire onward - give it to KIE • Assemble Video + Music as dialogue_audio and the music bed ducks under it, or to KIE • Save Video's chain via a mux step. Keep task_id for KIE • Reuse Completed Audio so a re-render doesn't cost twice, and log the whole thing with KIE • Save Generation Recipe.

    Chain the writing too: an LLM node (Gemini 3.8 Flash, or Claude Opus 5.5 for heavier work) can emit the dialogue_turns JSON, so your script and your cast stay in one graph.

    Install

    cd ComfyUI/custom_nodes
    git clone https://github.com/felipederosilva/ComfyUI-KIE-Nodes-Next
    python -m pip install -r ComfyUI-KIE-Nodes-Next/requirements.txt
    

    Then Settings → KIE.ai Nodes Next → Connection → Paste / replace KIE.ai API key. Paid node: every run submits a job.

    Where people get burned

    • Lite vs full. The pack also ships Gemini 3.8 Flash Lite TTS. The two nodes are identical apart from the model ID. Lite is the lighter tier - audition both on one real line and keep whichever reads your script better; don't assume the bigger name is automatically right for a short line.
    • Long scripts time out. Bump timeout_seconds past 1200 for anything substantial, or recover the finished job with the task_id rather than re-running it.
    • No word-level timings. You get a single audio file. Syncing dialogue to picture is your job, and it's the step that eats the afternoon.
    • JSON discipline. Both arrays must parse. Edit them in a text editor, not by hand-juggling quotes inside the widget.
    • It's hosted, so it filters. Content policy is Google's, and nothing in the node can route around it - which is the standing trade for every API node in ComfyUI.
    CategoryKIE Next/Audio/Google/Gemini TTS

    Inputs (7)

    NameTypeDefaultDescription
    speakersSTRING[ { "speaker_id": "Speaker 2", "voice_name": "Fenrir", "audio_profile": "ea Lorem sed", "accent": "Transatlantic", "style": "Whisper", "pace": "Natural" }, { "speaker_id": "Speaker 1", "voice_name": "Rasalgethi", "audio_profile": "ut Ut consectetur tempor", "accent": "British (Brixton)", "style": "Whisper", "pace": "Staccato" } ]List of speaker configurations
    dialogue_turnsSTRING[ { "speaker_id": "Speaker 1", "text": "Tero universe ad tremo utroque depulso valde molestiae antea. Certe deficio quidem cupio. Videlicet utrum sursum quas abduco conturbo currus conforto. Ustulo vereor abstergo adfectus arcus carpo subito careo. Voluntarius defluo virtus voluptatibus varius cornu tutamen canis vilis. Auctor adsuesco calcar bellicus dens aggredior aperio testimonium." } ]List of dialogue turns, output in sequential order
    temperatureoptFLOAT1.490–2Sampling temperature, e.g., 1
    sceneoptSTRINGminim mollit ullamco inScene description, e.g., "A quiet, warm room with a fireplace crackling softly."
    sample_contextoptSTRINGanim veniam Lorem voluptate exercitationSample context/overall tone, e.g., "Audiobook style narration. Tone is gentle and inviting."
    timeout_secondsoptINT120030–7200—
    callback_urloptSTRINGOptional KIE callback URL. Leave empty for normal ComfyUI use.

    Outputs (7)

    NameTypeDescription
    audioAUDIO—
    urlSTRING—
    all_urls_jsonSTRING—
    task_idSTRING—
    raw_jsonSTRING—
    credits_consumedFLOAT—
    credits_leftFLOAT—