ComfyUI Node

DramaBox TTS

DramaBox expressive TTS with voice cloning. Generates dramatic, expressive speech from a structured scene prompt. Requires ~24 GB VRAM on first run.

By kat3ri·Created 3 months ago·Updated 3 months ago· 19
DramaBox TTS
  • voice_sample
  • audio
textA woman speaks warmly, "Hello, how are you today?" She laughs, "Hahaha, it is so good to see you!"
cfg_scale2.5
stg_scale1.5
seed42
duration_multiplier1.10
Categoryaudio/DramaBox

Inputs (6)

NameTypeDefaultDescription
textSTRINGA woman speaks warmly, "Hello, how are you today?" She laughs, "Hahaha, it is so good to see you!"Scene prompt. Put dialogue in double quotes, stage directions outside them. Phonetic sounds (Hahaha, Hmm) go inside quotes; named actions (She sighs.) go outside.
cfg_scaleFLOAT2.51–10CFG guidance scale. Lower = more natural delivery; higher = more text-faithful. DramaBox default: 2.5.
stg_scaleFLOAT1.50–5Skip-token guidance scale. DramaBox default: 1.5.
voice_sampleoptAUDIOOptional voice reference for timbre cloning. 10+ seconds of clean speech recommended.
seedoptINT420–2147483647Random seed for reproducible generations.
duration_multiplieroptFLOAT1.100.5–3Multiply the auto-estimated speech duration. 1.1 adds 10 %% breathing room. Increase for slower delivery.

Outputs (1)

NameTypeDescription
audioAUDIO