Gemini 3.8 Flash Text to speech
Multi-speaker dialogue in one shot
- audio
- url
- all_urls_json
- task_id
- raw_json
- credits_consumed
- credits_left
What it is
A text-to-speech node that generates conversation, not narration: a cast of speakers defined once, a script of dialogue turns, and one AUDIO object back containing the whole exchange. It calls google/gemini-3-8-flash-tts on KIE, which relays Google's Gemini TTS.
The use case is specific and worth stating: you have a scene with two or three characters and you want their lines performed with distinct voices. Doing that with a single-voice TTS means stitching files and hand-matching tone. This does it in one job.
How it works
The widget shape is the whole node:
speakers- a JSON array of speaker configs. Each hasspeaker_id,voice_name(Gemini voices:Fenrir,Rasalgethi,Iapetus,Schedar…), and thenaudio_profile,accent,style,paceto describe delivery.dialogue_turns- a JSON array of{speaker_id, text}, performed in order.scene- the room and recording context, e.g. a quiet room with a fireplace crackling.sample_context- the overall tone and delivery reference, e.g. audiobook narration, gentle and inviting.temperature- 0 to 2. Tightens or loosens the performance.
The join key between the two arrays is speaker_id. Get it wrong and a line comes back in the wrong voice, or nothing comes back at all.
Under the hood it's an async KIE market task: submitted to /api/v1/jobs/createTask, polled until complete or until timeout_seconds (default 1200 s) elapses. Hence the task_id output, and hence the fact that a job can finish on KIE's side after your node has stopped waiting.
First thing to do: clear the placeholder defaults. The speakers and dialogue_turns values that ship in the widget are copied from KIE's docs example - two speakers with lorem-ipsum lines and a British (Brixton) accent. The node doesn't care that it's nonsense. It will synthesise it and charge you.
Inputs and outputs
Required: speakers, dialogue_turns. Optional: temperature, scene, sample_context, timeout_seconds, callback_url (leave empty in ComfyUI; it's for server-side callbacks).
Outputs: audio, url, all_urls_json, task_id, raw_json, credits_consumed, credits_left. The native audio output is what you wire onward - give it to KIE • Assemble Video + Music as dialogue_audio and the music bed ducks under it, or to KIE • Save Video's chain via a mux step. Keep task_id for KIE • Reuse Completed Audio so a re-render doesn't cost twice, and log the whole thing with KIE • Save Generation Recipe.
Chain the writing too: an LLM node (Gemini 3.8 Flash, or Claude Opus 5.5 for heavier work) can emit the dialogue_turns JSON, so your script and your cast stay in one graph.
Install
cd ComfyUI/custom_nodes
git clone https://github.com/felipederosilva/ComfyUI-KIE-Nodes-Next
python -m pip install -r ComfyUI-KIE-Nodes-Next/requirements.txt
Then Settings → KIE.ai Nodes Next → Connection → Paste / replace KIE.ai API key. Paid node: every run submits a job.
Where people get burned
- Lite vs full. The pack also ships
Gemini 3.8 Flash LiteTTS. The two nodes are identical apart from the model ID. Lite is the lighter tier - audition both on one real line and keep whichever reads your script better; don't assume the bigger name is automatically right for a short line. - Long scripts time out. Bump
timeout_secondspast 1200 for anything substantial, or recover the finished job with thetask_idrather than re-running it. - No word-level timings. You get a single audio file. Syncing dialogue to picture is your job, and it's the step that eats the afternoon.
- JSON discipline. Both arrays must parse. Edit them in a text editor, not by hand-juggling quotes inside the widget.
- It's hosted, so it filters. Content policy is Google's, and nothing in the node can route around it - which is the standing trade for every API node in ComfyUI.
Inputs (7)
| Name | Type | Default | Description |
|---|---|---|---|
| speakers | STRING | [ { "speaker_id": "Speaker 2", "voice_name": "Fenrir", "audio_profile": "ea Lorem sed", "accent": "Transatlantic", "style": "Whisper", "pace": "Natural" }, { "speaker_id": "Speaker 1", "voice_name": "Rasalgethi", "audio_profile": "ut Ut consectetur tempor", "accent": "British (Brixton)", "style": "Whisper", "pace": "Staccato" } ] | List of speaker configurations |
| dialogue_turns | STRING | [ { "speaker_id": "Speaker 1", "text": "Tero universe ad tremo utroque depulso valde molestiae antea. Certe deficio quidem cupio. Videlicet utrum sursum quas abduco conturbo currus conforto. Ustulo vereor abstergo adfectus arcus carpo subito careo. Voluntarius defluo virtus voluptatibus varius cornu tutamen canis vilis. Auctor adsuesco calcar bellicus dens aggredior aperio testimonium." } ] | List of dialogue turns, output in sequential order |
| temperatureopt | FLOAT | 1.490–2 | Sampling temperature, e.g., 1 |
| sceneopt | STRING | minim mollit ullamco in | Scene description, e.g., "A quiet, warm room with a fireplace crackling softly." |
| sample_contextopt | STRING | anim veniam Lorem voluptate exercitation | Sample context/overall tone, e.g., "Audiobook style narration. Tone is gentle and inviting." |
| timeout_secondsopt | INT | 120030–7200 | — |
| callback_urlopt | STRING | Optional KIE callback URL. Leave empty for normal ComfyUI use. |
Outputs (7)
| Name | Type | Description |
|---|---|---|
| audio | AUDIO | — |
| url | STRING | — |
| all_urls_json | STRING | — |
| task_id | STRING | — |
| raw_json | STRING | — |
| credits_consumed | FLOAT | — |
| credits_left | FLOAT | — |