ComfyUI Node
DramaBox TTS
DramaBox expressive TTS with voice cloning. Generates dramatic, expressive speech from a structured scene prompt. Requires ~24 GB VRAM on first run.
DramaBox TTS
- voice_sample
- audio
◄textA woman speaks warmly, "Hello, how are you today?" She laughs, "Hahaha, it is so good to see you!"►
◄cfg_scale2.5►
◄stg_scale1.5►
◄seed42►
◄duration_multiplier1.10►
Categoryaudio/DramaBox
Inputs (6)
| Name | Type | Default | Description |
|---|---|---|---|
| text | STRING | A woman speaks warmly, "Hello, how are you today?" She laughs, "Hahaha, it is so good to see you!" | Scene prompt. Put dialogue in double quotes, stage directions outside them. Phonetic sounds (Hahaha, Hmm) go inside quotes; named actions (She sighs.) go outside. |
| cfg_scale | FLOAT | 2.51–10 | CFG guidance scale. Lower = more natural delivery; higher = more text-faithful. DramaBox default: 2.5. |
| stg_scale | FLOAT | 1.50–5 | Skip-token guidance scale. DramaBox default: 1.5. |
| voice_sampleopt | AUDIO | Optional voice reference for timbre cloning. 10+ seconds of clean speech recommended. | |
| seedopt | INT | 420–2147483647 | Random seed for reproducible generations. |
| duration_multiplieropt | FLOAT | 1.100.5–3 | Multiply the auto-estimated speech duration. 1.1 adds 10 %% breathing room. Increase for slower delivery. |
Outputs (1)
| Name | Type | Description |
|---|---|---|
| audio | AUDIO | — |