ComfyUI Node
Maya1 TTS (AIO)
A ComfyUI node in audio/maya1 with 13 inputs and 1 output.
Maya1 TTS (AIO)
- audio
◄model_name(No models folder found - see console for instructions)►
◄dtypebfloat16►
◄attention_mechanismsdpa►
◄devicecuda►
◄voice_descriptionRealistic male voice in the 30s age with american accent. Normal pitch, warm timbre, conversational pacing.►
◄textHello! This is Maya1 <laugh> the best open source voice AI model with emotions.►
◄keep_model_in_vramtrue►
◄temperature0.40►
◄top_p0.90►
◄max_new_tokens4000►
◄repetition_penalty1.10►
◄seed0►
◄chunk_longformfalse►
Categoryaudio/maya1
Inputs (13)
| Name | Type | Default | Description |
|---|---|---|---|
| model_name | COMBO | (No models folder found - see console for instructions) | 1 options: (No models folder found - see console for instructions) |
| dtype | COMBO | bfloat16 | 5 options: 4bit (BNB), 8bit (BNB), float16, bfloat16, float32 |
| attention_mechanism | COMBO | sdpa | 4 options: sdpa, eager, flash_attention_2, sage_attention |
| device | COMBO | cuda | 2 options: cuda, cpu |
| voice_description | STRING | Realistic male voice in the 30s age with american accent. Normal pitch, warm timbre, conversational pacing. | — |
| text | STRING | Hello! This is Maya1 <laugh> the best open source voice AI model with emotions. | — |
| keep_model_in_vram | BOOLEAN | true | — |
| temperature | FLOAT | 0.400.1–2 | — |
| top_p | FLOAT | 0.900.1–1 | — |
| max_new_tokens | INT | 4000100–16000 | Maximum NEW SNAC tokens to generate per chunk (excludes input prompt tokens). Higher = longer audio per chunk (~50 tokens/word). 4000 tokens ≈ 30-40s audio |
| repetition_penalty | FLOAT | 1.101–2 | — |
| seed | INT | 00–18446744073709550000 | — |
| chunk_longform | BOOLEAN | false | Split long text into chunks at sentence boundaries with smooth crossfading. Enables unlimited audio length beyond the 18-20s limit |
Outputs (1)
| Name | Type | Description |
|---|---|---|
| audio | AUDIO | — |