Voice Design & Emotion Prompt Builder (Encyclopedic Qwen3/ChatterBox/Higgs)
Voice & Emotion Prompts for Qwen3-TTS, Chatterbox, and Higgs (Text Only)
- voice_design_prompt
- formatted_script
- engine_metadata
Let's get the name straight up front: "Voice Design & Emotion Prompt Builder" is a text node. It does not synthesize a single sound. It writes the prompt and the script you feed to an actual TTS model - which is still a real timesaver, because TTS engines all speak different dialects. Qwen3-TTS wants descriptive prose about the speaker; Chatterbox wants paralinguistic tags like [laughs] and [whisper]; Higgs Audio wants drama. Remembering all of that per project is a pain, and this node is built to forget it for you.
It's part of the ComfyUI-StudioPromptDirector pack, and it produces three outputs: voice_design_prompt (a descriptive prose brief for the voice), formatted_script (your dialogue with an emotion tag injected), and engine_metadata (a plain "Engine: X | Accent: Y | Timbre: Z" string you can pipe to a text display to remember what you generated).
The inputs
target_engine picks which of five engines you're writing for: Qwen3-TTS, ChatterBox, HiggsAudio v3, MOSS SoundEffect v2, or a "Universal TTS Suite" format. Then the voice knobs: vocal_gender_and_age, vocal_accent_origin (15 accents, from British RP to Egyptian English to Cockney), vocal_pitch_timbre, and emotional_delivery_preset - ten options from "Furious Screaming Outburst" to "Formal Neutral Broadcast". paralinguistic_tag_injector drops the actual bracket tag into your script. The Dialogue_Script_Text box holds the lines (defaults to the same Maya/Mariam argument from the rest of the pack), with Custom_Voice_Design_Prose and Acoustics_and_Environment_Notes as the fine-tuning overrides.
The author's README also sells "automatic prosody rules" - comma = 0.2s pause, period = 0.5s, double dash = 0.8s dramatic hesitation, CAPS = stress. Treat those as a house convention, not a model guarantee. No TTS engine advertises a fixed punctuation-to-timing table, so the rules are a sensible prompt-writing style, not something Chatterbox will honor to the millisecond.
The context you should know
The engines this targets are real and the community has opinions. Chatterbox (Resemble AI) is the open TTS that finally made people compare it to ElevenLabs, with an emotion-exaggeration dial and zero-shot cloning from ~5 seconds of audio. Higgs Audio is the multilingual heavyweight. One license warning if you're shipping anything: Higgs Audio v3 moved to a research/non-commercial license, a real regression from v2's Apache terms - check before you build a product on it. Also know that TTS inside ComfyUI is the fragile corner of the ecosystem: models arrive through hub packs like TTS Audio Suite, and dependency conflicts are the default failure mode. This node is nice precisely because it's pure Python and adds no torch/transformers risk to that pile.
Install
Same as the rest of the pack - nothing to download:
cd ComfyUI/custom_nodes
git clone https://github.com/nexusfinancial-dev/ComfyUI-StudioPromptDirector.git
Or search "ComfyUI-StudioPromptDirector" in ComfyUI Manager and restart.
Wiring
voice_design_prompt goes to whatever voice-designer input your engine exposes (Qwen3-TTS style), formatted_script to the script/speech input, and engine_metadata to a text display if you want provenance. If an engine mangles your output, check the mismatch first: if you picked Qwen3-TTS but the script is carrying [shouting] brackets, that's a Chatterbox-style tag doing nothing for a prose-driven engine - pick the engine first, then let the tag injector match it. And keep the target_engine honest: "Universal" is a convenience format, not a magic bullet that makes every TTS understand every tag.
Inputs (9)
| Name | Type | Default | Description |
|---|---|---|---|
| target_engine | COMBO | Qwen3-TTS (Textual Voice Design) | 5 options: Qwen3-TTS (Textual Voice Design), ChatterBox (Paralinguistic Tags & Prosody Rules), HiggsAudio v3 (Dramatic Multi-Speaker Engine), MOSS SoundEffect v2 (Background Ambience & SFX), TTS Audio Suite Universal Format |
| vocal_gender_and_age | COMBO | Young Adult Female (Early 20s - Bright, Crisp & Melodic) | 9 options: Young Adult Female (Early 20s - Bright, Crisp & Melodic), Adult Female (Mid 20s-30s - Warm, Articulate & Professional), Mature Female (40s-50s - Deep, Rich & Authoritative), Elderly Female (60+ Years - Raspy, Gentle & Wise), Young Adult Male (Early 20s - Casual, Energetic & Confident), Adult Male (30s-40s - Deep Resonant Baritone), +3 |
| vocal_accent_origin | COMBO | American Modern Standard (Clean Studio Accent) | 15 options: American Modern Standard (Clean Studio Accent), Egyptian / Middle Eastern English (Warm Melodic Cadence), British Received Pronunciation (Elegant, Articulate & Refined), Cockney London Working-Class Accent, Scottish Resonant Celtic Accent, Irish Melodic Lilt Accent, +9 |
| vocal_pitch_timbre | COMBO | Medium-High Pitch (Bright, Clear & Energetic) | 8 options: Medium-High Pitch (Bright, Clear & Energetic), Warm Mezzo-Soprano (Gentle, Melodic & Soothing), Deep Resonant Baritone (Authoritative, Grounded & Heavy), Husky Breathy Tone (Intimate, Vulnerable & Emotional), Raspy Gravelly Timbre (Intense, Agitated & Weathered), Silky Smooth Velvet Tone (Commercial & Luxurious), +2 |
| emotional_delivery_preset | COMBO | Furious Screaming Outburst [shouting] | 10 options: Furious Screaming Outburst [shouting], Shocked Incredulous Gasp & Trembling Voice [gasps], Calm Warm Conversational Chuckle [laughs], Tense Suppressed Anger & Controlled Cold Whispers [whisper], Deep Heartbroken Sadness & Tearful Sobbing [crying], Sarcastic Mocking Smirk & Chuckle [snicker], +4 |
| paralinguistic_tag_injector | COMBO | [shouting] - Forced Volume & Acoustic Stress | 11 options: [shouting] - Forced Volume & Acoustic Stress, [gasps] - Sharp Inhalation of Shock, [laughs] - Natural Conversational Chuckle, [crying] - Tearful Break in Voice, [sighs] - Deep Exhalation of Exhaustion, [whisper] - Quiet Intimate Tone, +5 |
| 🏷️_Dialogue_Script_Textopt | STRING | Maya: I told you never to speak to me like that! Mariam: How dare you slap me! | — |
| 🏷️_Custom_Voice_Design_Proseopt | STRING | — | |
| 🏷️_Acoustics_and_Environment_Notesopt | STRING | pristine studio microphone acoustics, zero room echo, warm proximity effect | — |
Outputs (3)
| Name | Type | Description |
|---|---|---|
| voice_design_prompt | STRING | — |
| formatted_script | STRING | — |
| engine_metadata | STRING | — |