Nodes/ComfyUI-Gemini_3x_Pro/πŸ”Š Gemini Text-to-Speech v2
ComfyUI Node

πŸ”Š Gemini Text-to-Speech v2

It's a Prompt, Not a Slider

By asirusasr-makerΒ·Created 2 months agoΒ·Updated a day agoΒ· 6
πŸ”Š Gemini Text-to-Speech v2
    • audio
    • tts_info
    β—„textHello, this is a Gemini 3.8 TTS Pro test.β–Ί
    β—„modelgemini-3.8-flash-ttsβ–Ί
    β—„voiceKoreβ–Ί
    β—„languageAuto Detectβ–Ί
    β—„emotionNeutralβ–Ί
    β—„style_presetNaturalβ–Ί
    β—„narration_modeNaturalβ–Ί
    β—„speaker_profileNoneβ–Ί
    β—„paceNaturalβ–Ί
    β—„pitch_modeNaturalβ–Ί
    β—„custom_voice_idβ–Ί
    β—„custom_languageβ–Ί
    β—„accentAutoβ–Ί
    β—„custom_accentβ–Ί
    β—„custom_emotionβ–Ί
    β—„custom_styleβ–Ί
    β—„custom_narrationβ–Ί
    β—„custom_profileβ–Ί
    β—„inline_vocal_tagstrueβ–Ί
    β—„custom_paceβ–Ί
    β—„custom_pitchβ–Ί
    β—„speed1.00β–Ί
    β—„pitch0.0β–Ί
    β—„api_keyβ–Ί
    β—„proxyβ–Ί
    β—„fallback_enabledtrueβ–Ί
    β—„retries_per_model1β–Ί
    β—„cooldown_seconds30β–Ί

    Why you'd use a hosted voice at all

    Local TTS is genuinely good now. Chatterbox clones a voice from five seconds of audio at a quality people stopped complaining about, Kokoro runs on a CPU, and neither costs anything per line. So the honest case for this node is narrower than the marketing: you want a competent voice right now, with no model download, no transformers version roulette, and no 4 GB of weights sitting next to your checkpoint.

    That last part matters more than it sounds. Audio in ComfyUI is still a retrofitted layer - every local TTS option arrives as a node pack with its own dependency stack, and the maintainers of those packs will tell you the default failure mode is one engine's update breaking three others. This node has no local model. Its requirements are the SDK and pillow. If you have a graph that needs a line of narration and you don't want to touch your Python environment, that's the trade you're making.

    Where it doesn't win: a fixed narrator across dozens of clips. Per-call pricing on dozens of clips adds up, and a local model you ran once is free forever after.

    How it works

    You send text plus delivery instructions, and the model returns audio. The node asks for 24 kHz output and decodes it deterministically - raw L16/PCM goes through the PCM path at its declared rate, and WAV payloads are detected by their RIFF header and read with the standard wave reader, so both current and older response formats land in the same place.

    The bit that trips people up: the delivery controls are prose. speed and pitch are legacy compatibility inputs kept so older workflows don't break, and the node translates them into phrases like "slower overall pacing" or "slightly lower pitch" which it hands the model as style guidance. There is no numeric rate parameter on the far end. A speed of 1.4 is a strong suggestion, not 1.4Γ—.

    Same story for style, which in this page's schema is a free-text delivery instruction with a sensible default of "Natural, clear, expressive narration." Treat it as a one-line stage direction - "warm and unhurried, slight breathiness, news-anchor clarity" - rather than a preset. Later builds of the pack (v2.0.6) turned this area into a much larger set of structured delivery presets, so if your node shows more fields than this page lists, that's why; the changelog in the repo is the current truth.

    Inputs and outputs

    Required: text, model (gemini-3.8-flash-tts or gemini-3.8-flash-lite-tts), voice. The eight voices here - Puck, Charon, Kore, Fenrir, Leda, Orus, Aoede, Callirhoe - are the standard Gemini prebuilt set, so they'll be familiar if you've used Gemini speech anywhere else. Kore is the default and it's a solid neutral narrator.

    Optional: style, speed, pitch, api_key, proxy, fallback_enabled, retries_per_model, cooldown_seconds. Fallback runs Flash TTS down to Flash-Lite TTS when the bigger model is unavailable.

    Outputs are audio (ComfyUI AUDIO - straight into Save Audio, or the audio input of a lip-sync workflow) and tts_info (a string of metadata about the run).

    Install

    Manager, searching the display name ComfyUI Gemini 3x Pro - or clone it by hand, which is what the pack's README documents:

    cd ComfyUI/custom_nodes
    git clone https://github.com/asirusasr-maker/ComfyUI-Gemini_3x_Pro
    

    Then the dependencies, with the interpreter that runs your ComfyUI:

    python_embeded\python.exe -m pip install -r ComfyUI\custom_nodes\ComfyUI-Gemini_3x_Pro\requirements.txt
    

    That's google-genai, pillow, numpy and sounddevice - no audio model, no torchaudio requirement from this pack. Set GEMINI_API_KEY in the pack's config.json or as an environment variable (the node's api_key field wins over both, but remember it gets written into any workflow JSON you share), then restart.

    Where people get burned

    One execution, one request. There's no chunking, so a 2,000-word script is a single call with a single response - long, billed as one job, and much likelier to hit a transient failure partway through a nice take. Split long scripts into paragraphs, generate them as separate runs, and concatenate. You'll get better pacing control that way too, since each paragraph can carry its own style.

    Long text and 429s go together. The retry router backs off and falls back to the Lite model, which is the right behaviour but also means a busy period can quietly hand you Lite-quality audio. tts_info is where you check which model actually spoke.

    There's no 401-proofing to do. Auth errors surface immediately and are never retried. If every run fails instantly, that's a key or config problem, not a rate limit - and note the pack refuses to treat the placeholder string your_api_key_here as a real key, so a half-filled config.json behaves like no key at all.

    Wrong sample rate assumptions in downstream nodes. The output lands as an AUDIO object, so most consumers are fine, but anything expecting 44.1 kHz specifically wants a resample in between. This is the one audio node in the pack you can use without a microphone, an ffmpeg build, or a GPU.

    CategoryGemini 3.x

    Inputs (28)

    NameTypeDefaultDescription
    textSTRINGHello, this is a Gemini 3.8 TTS Pro test.β€”
    modelCOMBOgemini-3.8-flash-tts2 options: gemini-3.8-flash-tts, gemini-3.8-flash-lite-tts
    voiceCOMBOKore30 options: Zephyr, Puck, Charon, Kore, Fenrir, Leda, +24
    languageCOMBOAuto Detect24 options: Auto Detect, English, Russian, Uzbek, Spanish, French, +18
    emotionCOMBONeutral20 options: Neutral, Calm, Warm, Friendly, Happy, Excited, +14
    style_presetCOMBONatural16 options: Natural, Conversational, Narration, Documentary, Cinematic, Audiobook, +10
    narration_modeCOMBONatural12 options: Natural, Narrator, Documentary, News Anchor, YouTube Creator, Audiobook, +6
    speaker_profileCOMBONone11 options: None, Documentary Male, Documentary Female, YouTube Narrator, Luxury Commercial, News Presenter, +5
    paceCOMBONatural7 options: Natural, Very Slow, Slow, Relaxed, Fast, Very Fast, +1
    pitch_modeCOMBONatural7 options: Natural, Low, Medium, High, Very High, Monotone, +1
    custom_voice_idoptSTRINGOptional Extended Voice Library, Voice Design (voice_...) or Voice Replication ID. Overrides the prebuilt Voice.
    custom_languageoptSTRINGUsed when Language is Other / Custom.
    accentoptCOMBOAuto14 options: Auto, American English, British English, Australian English, Indian English, Irish English, +8
    custom_accentoptSTRINGUsed when Accent is Custom.
    custom_emotionoptSTRINGUsed when Emotion is Custom.
    custom_styleoptSTRINGExtra turn-level delivery instructions sent through speech_metadata.style.
    custom_narrationoptSTRINGAdditional narration direction.
    custom_profileoptSTRINGCustom reusable speaker persona/delivery description.
    inline_vocal_tagsoptBOOLEANtrueKeep supported <laugh>, <sigh>, <breath>, <short pause>, etc. in the transcript.
    custom_paceoptSTRINGUsed when Pace is Custom, e.g. 'speaking at a relaxed pace'.
    custom_pitchoptSTRINGUsed when Pitch is Custom, e.g. 'slightly low pitch with warm inflection'.
    speedoptFLOAT1.000.5–2Legacy compatibility control. Converted to natural-language pace guidance.
    pitchoptFLOAT0.0-10–10Legacy compatibility control. Converted to natural-language pitch guidance.
    api_keyoptSTRINGβ€”
    proxyoptSTRINGβ€”
    fallback_enabledoptBOOLEANtrueβ€”
    retries_per_modeloptINT10–4β€”
    cooldown_secondsoptFLOAT300–300β€”

    Outputs (2)

    NameTypeDescription
    audioAUDIOβ€”
    tts_infoSTRINGβ€”