Gemini 3.8 Flash Lite Text to speech
A whole scene of dialogue, not one voice
- audio
- url
- all_urls_json
- task_id
- raw_json
- credits_consumed
- credits_left
What it is
Most TTS nodes take one string and give you one voice. This one takes a cast list and a script - several speakers, each with a voice, an accent, a style and a pace - and a list of dialogue turns, and returns a single AUDIO of the conversation.
If you're building a short film in ComfyUI, that's a different thing entirely from narration. You get multi-speaker dialogue as one artefact, at the length your scene actually is, ready to feed into the audio side of an assembly step.
How it works
This node is a KIE market task, not a direct API call: it posts google/gemini-3-8-flash-lite-tts to /api/v1/jobs/createTask, gets a task ID, and polls until the job lands or timeout_seconds runs out. That's why there's a timeout widget and a task_id output - it's asynchronous work on KIE's side, and it can take a while for a long script.
Two JSON widgets carry the actual content:
speakers- an array of speaker configs. Each entry has aspeaker_id, avoice_name(Gemini voice names likeIapetus,Schedar,Fenrir,Rasalgethi), plusaudio_profile,accent,styleandpacedescribing the delivery.dialogue_turns- an array of{speaker_id, text}objects, spoken in array order.
The speaker_id values are the join between the two lists, so a typo there is how you get silence or the wrong voice on a line.
Around those sit two knobs that do more than they look like they do: scene ("A quiet, warm room with a fireplace crackling softly.") sets the room and the recording character, and sample_context sets overall tone ("Audiobook style narration. Tone is gentle and inviting."). temperature (0–2, default around 0.3 on this node) loosens or tightens delivery.
Read the defaults before you run it. The shipped speakers and dialogue_turns values are the placeholder junk from KIE's own docs example - lorem-ipsum dialogue, an accent of "American (Valley)", a pace of "Rapid Fire". Delete all of it. This trips people up constantly, because the node will happily synthesise nonsense and bill you for it.
Inputs and outputs
Both JSON widgets are required, plus the optional temperature, scene, sample_context, timeout_seconds (default 1200 s) and callback_url (leave empty - it's for server-side workflows, not Comfy graphs).
Outputs: audio (the native AUDIO object), url, all_urls_json, task_id, raw_json, credits_consumed, credits_left. Keep the task_id: paste it into KIE • Reuse Completed Audio later to re-download the same result without paying again, and into KIE • Save Generation Recipe to record what you made.
Feed the audio straight into KIE • Assemble Video + Music's dialogue_audio input and the music ducks under it automatically.
Install
cd ComfyUI/custom_nodes
git clone https://github.com/felipederosilva/ComfyUI-KIE-Nodes-Next
python -m pip install -r ComfyUI-KIE-Nodes-Next/requirements.txt
Then set your key once in Settings → KIE.ai Nodes Next → Connection. This node is a paid generator - every execution submits a real job and spends credits, so audition it with two lines, not your whole script.
Where people get burned
- Valid JSON or nothing. Both arrays are parsed locally before upload and a malformed array is a hard error. A trailing comma in a hand-edited cast list is the classic.
- Lite vs full. The pack ships two otherwise identical nodes for this family -
liteand the plain Flash TTS. The schemas are the same and the only difference is the model ID behind them. Draft on one, keep the take from the other. - Timeout, not failure. A long script can exceed
timeout_secondsand the node will give up waiting even though KIE finished. Take thetask_id, and useKIE • Reuse Completed Audioto collect the result instead of running the job again. - No captions, no timing. You get audio, not word timings. If you need to line dialogue up to picture, that alignment step is on you.
Inputs (7)
| Name | Type | Default | Description |
|---|---|---|---|
| speakers | STRING | [ { "speaker_id": "Speaker 1", "voice_name": "Iapetus", "audio_profile": "exercitation", "accent": "American (Valley)", "style": "Newscaster", "pace": "Rapid Fire" }, { "speaker_id": "Speaker 2", "voice_name": "Schedar", "audio_profile": "ex Lorem ut non", "accent": "British (RP)", "style": "Whisper", "pace": "The Drift" } ] | List of speaker configurations |
| dialogue_turns | STRING | [ { "speaker_id": "Speaker 1", "text": "Varietas votum cerno aspernatur communis utrum tempora apostolus usque demitto. Supra charisma desipio. Bellicus deorsum speciosus coepi excepturi tres timor. Attollo caecus spargo repudiandae cohors occaecati vespillo vulgaris cometes. Synagoga capillus verus universe. Canis una uxor annus desino causa templum." } ] | List of dialogue turns, output in sequential order |
| temperatureopt | FLOAT | 0.300–2 | Sampling temperature, e.g., 1 |
| sceneopt | STRING | ut | Scene description, e.g., "A quiet, warm room with a fireplace crackling softly." |
| sample_contextopt | STRING | ipsum minim | Sample context/overall tone, e.g., "Audiobook style narration. Tone is gentle and inviting." |
| timeout_secondsopt | INT | 120030–7200 | — |
| callback_urlopt | STRING | Optional KIE callback URL. Leave empty for normal ComfyUI use. |
Outputs (7)
| Name | Type | Description |
|---|---|---|
| audio | AUDIO | — |
| url | STRING | — |
| all_urls_json | STRING | — |
| task_id | STRING | — |
| raw_json | STRING | — |
| credits_consumed | FLOAT | — |
| credits_left | FLOAT | — |