ComfyUI Node
Dia text to speech
A ComfyUI node in audio/dia with 14 inputs and 1 output.
Dia text to speech
- input_audio
- audio
◄model_pathmodels/Dia/dia-v0_1.pth►
◄seed12345►
◄save_audio_filetrue►
◄filename_prefixaudio/dia►
◄speech[S1] Dia is an open weights text to dialogue model.
[S2] You get full control over scripts and voices.
[S1] Wow. Amazing. (laughs)
[S2] Try it now on Git hub or Hugging Face.►
◄cfg_scale3.00►
◄temperature1.30►
◄top_p0.95►
◄use_cfg_filtertrue►
◄use_torch_compilefalse►
◄cfg_filter_top_k35►
◄input_audio_transcript►
◄available_tags(laughs), (clears throat), (sighs), (gasps), (coughs),
(singing), (sings), (mumbles), (beep), (groans),
(sniffs), (claps), (screams), (inhales), (exhales),
(applause), (burps), (humming), (sneezes), (chuckle), (whistles)►
Categoryaudio/dia
Inputs (14)
| Name | Type | Default | Description |
|---|---|---|---|
| model_path | STRING | models/Dia/dia-v0_1.pth | — |
| seed | INT | 123450–2147483647 | — |
| save_audio_file | BOOLEAN | true | — |
| filename_prefix | STRING | audio/dia | — |
| speech | STRING | [S1] Dia is an open weights text to dialogue model. [S2] You get full control over scripts and voices. [S1] Wow. Amazing. (laughs) [S2] Try it now on Git hub or Hugging Face. | — |
| cfg_scale | FLOAT | 3.000–10 | — |
| temperature | FLOAT | 1.300–10 | — |
| top_p | FLOAT | 0.950–10 | — |
| use_cfg_filter | BOOLEAN | true | — |
| use_torch_compile | BOOLEAN | false | — |
| cfg_filter_top_k | INT | 350–100 | — |
| input_audioopt | AUDIO | — | |
| input_audio_transcriptopt | STRING | — | |
| available_tagsopt | STRING | (laughs), (clears throat), (sighs), (gasps), (coughs), (singing), (sings), (mumbles), (beep), (groans), (sniffs), (claps), (screams), (inhales), (exhales), (applause), (burps), (humming), (sneezes), (chuckle), (whistles) | — |
Outputs (1)
| Name | Type | Description |
|---|---|---|
| audio | AUDIO | — |