ComfyUI Node
Gemini Text-to-Speech
A ComfyUI node in API Toolkit/Gemini/Audio with 6 inputs and 1 output.
Gemini Text-to-Speech
- audio
◄api_key►
◄modelgemini-2.5-flash-preview-tts►
◄text►
◄voiceKore►
◄custom_model►
◄style_prompt►
CategoryAPI Toolkit/Gemini/Audio
Inputs (6)
| Name | Type | Default | Description |
|---|---|---|---|
| api_key | STRING | Gemini API key. Leave blank to use GEMINI_API_KEY env var. | |
| model | COMBO | gemini-2.5-flash-preview-tts | TTS model. Pro = higher quality, Flash = faster. |
| text | STRING | Text to speak. Can include speaker tags for multi-speaker dialogue. | |
| voice | COMBO | Kore | Prebuilt voice. Each has different characteristics. |
| custom_modelopt | STRING | — | |
| style_promptopt | STRING | Optional style instruction prepended to text (e.g., 'Say cheerfully:'). |
Outputs (1)
| Name | Type | Description |
|---|---|---|
| audio | AUDIO | — |