π Gemini Text-to-Speech v2
It's a Prompt, Not a Slider
- audio
- tts_info
Why you'd use a hosted voice at all
Local TTS is genuinely good now. Chatterbox clones a voice from five seconds of audio at a quality people stopped complaining about, Kokoro runs on a CPU, and neither costs anything per line. So the honest case for this node is narrower than the marketing: you want a competent voice right now, with no model download, no transformers version roulette, and no 4 GB of weights sitting next to your checkpoint.
That last part matters more than it sounds. Audio in ComfyUI is still a retrofitted layer - every local TTS option arrives as a node pack with its own dependency stack, and the maintainers of those packs will tell you the default failure mode is one engine's update breaking three others. This node has no local model. Its requirements are the SDK and pillow. If you have a graph that needs a line of narration and you don't want to touch your Python environment, that's the trade you're making.
Where it doesn't win: a fixed narrator across dozens of clips. Per-call pricing on dozens of clips adds up, and a local model you ran once is free forever after.
How it works
You send text plus delivery instructions, and the model returns audio. The node asks for 24 kHz output and decodes it deterministically - raw L16/PCM goes through the PCM path at its declared rate, and WAV payloads are detected by their RIFF header and read with the standard wave reader, so both current and older response formats land in the same place.
The bit that trips people up: the delivery controls are prose. speed and pitch are legacy compatibility inputs kept so older workflows don't break, and the node translates them into phrases like "slower overall pacing" or "slightly lower pitch" which it hands the model as style guidance. There is no numeric rate parameter on the far end. A speed of 1.4 is a strong suggestion, not 1.4Γ.
Same story for style, which in this page's schema is a free-text delivery instruction with a sensible default of "Natural, clear, expressive narration." Treat it as a one-line stage direction - "warm and unhurried, slight breathiness, news-anchor clarity" - rather than a preset. Later builds of the pack (v2.0.6) turned this area into a much larger set of structured delivery presets, so if your node shows more fields than this page lists, that's why; the changelog in the repo is the current truth.
Inputs and outputs
Required: text, model (gemini-3.8-flash-tts or gemini-3.8-flash-lite-tts), voice. The eight voices here - Puck, Charon, Kore, Fenrir, Leda, Orus, Aoede, Callirhoe - are the standard Gemini prebuilt set, so they'll be familiar if you've used Gemini speech anywhere else. Kore is the default and it's a solid neutral narrator.
Optional: style, speed, pitch, api_key, proxy, fallback_enabled, retries_per_model, cooldown_seconds. Fallback runs Flash TTS down to Flash-Lite TTS when the bigger model is unavailable.
Outputs are audio (ComfyUI AUDIO - straight into Save Audio, or the audio input of a lip-sync workflow) and tts_info (a string of metadata about the run).
Install
Manager, searching the display name ComfyUI Gemini 3x Pro - or clone it by hand, which is what the pack's README documents:
cd ComfyUI/custom_nodes
git clone https://github.com/asirusasr-maker/ComfyUI-Gemini_3x_Pro
Then the dependencies, with the interpreter that runs your ComfyUI:
python_embeded\python.exe -m pip install -r ComfyUI\custom_nodes\ComfyUI-Gemini_3x_Pro\requirements.txt
That's google-genai, pillow, numpy and sounddevice - no audio model, no torchaudio requirement from this pack. Set GEMINI_API_KEY in the pack's config.json or as an environment variable (the node's api_key field wins over both, but remember it gets written into any workflow JSON you share), then restart.
Where people get burned
One execution, one request. There's no chunking, so a 2,000-word script is a single call with a single response - long, billed as one job, and much likelier to hit a transient failure partway through a nice take. Split long scripts into paragraphs, generate them as separate runs, and concatenate. You'll get better pacing control that way too, since each paragraph can carry its own style.
Long text and 429s go together. The retry router backs off and falls back to the Lite model, which is the right behaviour but also means a busy period can quietly hand you Lite-quality audio. tts_info is where you check which model actually spoke.
There's no 401-proofing to do. Auth errors surface immediately and are never retried. If every run fails instantly, that's a key or config problem, not a rate limit - and note the pack refuses to treat the placeholder string your_api_key_here as a real key, so a half-filled config.json behaves like no key at all.
Wrong sample rate assumptions in downstream nodes. The output lands as an AUDIO object, so most consumers are fine, but anything expecting 44.1 kHz specifically wants a resample in between. This is the one audio node in the pack you can use without a microphone, an ffmpeg build, or a GPU.
Inputs (28)
| Name | Type | Default | Description |
|---|---|---|---|
| text | STRING | Hello, this is a Gemini 3.8 TTS Pro test. | β |
| model | COMBO | gemini-3.8-flash-tts | 2 options: gemini-3.8-flash-tts, gemini-3.8-flash-lite-tts |
| voice | COMBO | Kore | 30 options: Zephyr, Puck, Charon, Kore, Fenrir, Leda, +24 |
| language | COMBO | Auto Detect | 24 options: Auto Detect, English, Russian, Uzbek, Spanish, French, +18 |
| emotion | COMBO | Neutral | 20 options: Neutral, Calm, Warm, Friendly, Happy, Excited, +14 |
| style_preset | COMBO | Natural | 16 options: Natural, Conversational, Narration, Documentary, Cinematic, Audiobook, +10 |
| narration_mode | COMBO | Natural | 12 options: Natural, Narrator, Documentary, News Anchor, YouTube Creator, Audiobook, +6 |
| speaker_profile | COMBO | None | 11 options: None, Documentary Male, Documentary Female, YouTube Narrator, Luxury Commercial, News Presenter, +5 |
| pace | COMBO | Natural | 7 options: Natural, Very Slow, Slow, Relaxed, Fast, Very Fast, +1 |
| pitch_mode | COMBO | Natural | 7 options: Natural, Low, Medium, High, Very High, Monotone, +1 |
| custom_voice_idopt | STRING | Optional Extended Voice Library, Voice Design (voice_...) or Voice Replication ID. Overrides the prebuilt Voice. | |
| custom_languageopt | STRING | Used when Language is Other / Custom. | |
| accentopt | COMBO | Auto | 14 options: Auto, American English, British English, Australian English, Indian English, Irish English, +8 |
| custom_accentopt | STRING | Used when Accent is Custom. | |
| custom_emotionopt | STRING | Used when Emotion is Custom. | |
| custom_styleopt | STRING | Extra turn-level delivery instructions sent through speech_metadata.style. | |
| custom_narrationopt | STRING | Additional narration direction. | |
| custom_profileopt | STRING | Custom reusable speaker persona/delivery description. | |
| inline_vocal_tagsopt | BOOLEAN | true | Keep supported <laugh>, <sigh>, <breath>, <short pause>, etc. in the transcript. |
| custom_paceopt | STRING | Used when Pace is Custom, e.g. 'speaking at a relaxed pace'. | |
| custom_pitchopt | STRING | Used when Pitch is Custom, e.g. 'slightly low pitch with warm inflection'. | |
| speedopt | FLOAT | 1.000.5β2 | Legacy compatibility control. Converted to natural-language pace guidance. |
| pitchopt | FLOAT | 0.0-10β10 | Legacy compatibility control. Converted to natural-language pitch guidance. |
| api_keyopt | STRING | β | |
| proxyopt | STRING | β | |
| fallback_enabledopt | BOOLEAN | true | β |
| retries_per_modelopt | INT | 10β4 | β |
| cooldown_secondsopt | FLOAT | 300β300 | β |
Outputs (2)
| Name | Type | Description |
|---|---|---|
| audio | AUDIO | β |
| tts_info | STRING | β |