ComfyUI Node
Scenema Audio Generate ⚡
Expressive text-to-speech with zero-shot voice cloning on the Scenema audio diffusion model.
Scenema Audio Generate ⚡
- model
- clip
- vae
- ref_latent
- identity_reference
- audio
◄presetCustom►
◄voice_descriptionMale, late 60s. Deep, gravelly. Slow and deliberate. The weight of the cosmos in every word.►
◄gender▾►
◄speech_textLook again at that dot. That's here. That's home. That's us.►
◄scene▾►
◄seed0►
◄custom_scene►
◄action_tags►
◄languageEnglish►
◄pace1.5►
◄skip_vcfalse►
◄vc_steps25►
◄vc_cfg_rate0.50►
Category🤖 CCTech/Scenema
Inputs (18)
| Name | Type | Default | Description |
|---|---|---|---|
| model | MODEL | — | |
| clip | CLIP | — | |
| vae | VAE | — | |
| preset | COMBO | Custom | Apply a preset's voice, gender, scene, and performance tags. The speech_text field is always used as entered. |
| voice_description | STRING | Male, late 60s. Deep, gravelly. Slow and deliberate. The weight of the cosmos in every word. | Describe the voice: age, gender presentation, timbre, accent, delivery style. |
| gender | COMBO | Grammatical gender used for pronouns in the compiled prompt. | |
| speech_text | STRING | Look again at that dot. That's here. That's home. That's us. | The text to speak. Use [bracketed cues] inline for mid-speech performance direction: [He laughs], [She whispers]. Long text is auto-split at sentence boundaries. |
| scene | COMBO | Acoustic environment injected into the prompt. | |
| seed | INT | 00–18446744073709550000 | — |
| custom_sceneopt | STRING | Freeform scene description. When non-empty, overrides the scene dropdown. | |
| action_tagsopt | STRING | Delivery cues, one per line. Each becomes a stage direction the model performs. | |
| languageopt | COMBO | English | Target language. Write speech_text in that language. |
| paceopt | FLOAT | 1.50.5–3 | Duration budget multiplier. Higher = slower speech. 1.5 is the validated default. |
| ref_latentopt | LATENT | Optional voice reference (Scenema VAE Encode output) for zero-shot cloning. | |
| skip_vcopt | BOOLEAN | false | Skip the final SeedVC identity-consistency pass. |
| identity_referenceopt | AUDIO | Optional fixed voice identity for SeedVC. This is separate from the LTX A2V ref_latent input. | |
| vc_stepsopt | INT | 251–200 | — |
| vc_cfg_rateopt | FLOAT | 0.500–2 | — |
Outputs (1)
| Name | Type | Description |
|---|---|---|
| audio | AUDIO | — |