ComfyUI Node
arkennemasis fal · MiniMax Speech 2.8 HD · $0.1/1k chars
Generate speech from text prompts and different voices using the MiniMax Speech-2.8 HD model, which leverages advanced AI techniques to create high-quality text-to-speech. Pricing (from fal, 2026-09-29): $0.1 per 1000 character (fal's billing unit). Model page: https://fal.ai/models/fal-ai/minimax/speech-2.8-hd Key: FAL_KEY in the .env next to run_nvidia_gpu.bat. Results are saved in output/fal/fal-ai_minimax_speech-2.8-hd/.
arkennemasis fal · MiniMax Speech 2.8 HD · $0.1/1k chars
- audio
- audio_path
- duration_ms
- info
◄prompt►
◄voice_setting_speed1.00►
◄voice_setting_vol1.00►
◄voice_setting_english_normalizationfalse►
◄voice_setting_pitch0►
◄voice_setting_emotion(not set)►
◄voice_setting_voice_idWise_Woman►
◄audio_setting_formatmp3►
◄audio_setting_bitrate128000►
◄audio_setting_sample_rate32000►
◄audio_setting_channel1►
◄language_boost(not set)►
◄output_formathex►
◄pronunciation_dict►
◄normalization_setting_target_range8.00►
◄normalization_setting_target_loudness-18.00►
◄normalization_setting_enabledtrue►
◄normalization_setting_target_peak-0.50►
◄voice_modify_pitch0►
◄voice_modify_intensity0►
◄voice_modify_timbre0►
◄max_cost_usd20.0►
◄reuse_identical_runtrue►
◄extra_json►
◄max_concurrent1►
Categoryarkennemasis/fal/Audio/MiniMax
Inputs (25)
| Name | Type | Default | Description |
|---|---|---|---|
| prompt | STRING | Text to convert to speech. Use `<#x#>` for pauses (x = 0.01-99.99 seconds). Supports interjection tags: `(laughs)`, `(sighs)`, `(coughs)`, `(clears throat)`, `(gasps)`, `(sniffs)`, `(groans)`, `(yawns)`. | |
| voice_setting_speed | FLOAT | 1.000.5–2 | Speech speed (0.5-2.0) |
| voice_setting_vol | FLOAT | 1.000.01–10 | Volume (0-10) |
| voice_setting_english_normalization | BOOLEAN | false | Enables English text normalization to improve number reading performance, with a slight increase in latency |
| voice_setting_pitch | INT | 0-12–12 | Voice pitch (-12 to 12) |
| voice_setting_emotion | COMBO | (not set) | Emotion of the generated speech '(not set)' leaves it to the model. |
| voice_setting_voice_id | STRING | Wise_Woman | Predefined voice ID to use for synthesis Empty leaves it out. |
| audio_setting_format | COMBO | mp3 | Audio format |
| audio_setting_bitrate | COMBO | 128000 | Bitrate of generated audio |
| audio_setting_sample_rate | COMBO | 32000 | Sample rate of generated audio |
| audio_setting_channel | COMBO | 1 | Number of audio channels (1=mono, 2=stereo) |
| language_boost | COMBO | (not set) | Enhance recognition of specified languages and dialects '(not set)' leaves it to the model. |
| output_format | COMBO | hex | Format of the output content (non-streaming only) |
| pronunciation_dict | STRING | Custom pronunciation dictionary for text replacement JSON. Empty leaves it out. | |
| normalization_setting_target_range | FLOAT | 8.000–20 | Target loudness range in LU (default 8.0) |
| normalization_setting_target_loudness | FLOAT | -18.00-70–-10 | Target loudness in LUFS (default -18.0) |
| normalization_setting_enabled | BOOLEAN | true | Enable loudness normalization for the audio |
| normalization_setting_target_peak | FLOAT | -0.50-3–0 | Target peak level in dBTP (default -0.5). |
| voice_modify_pitch | INT | 0-100–100 | Pitch adjustment in semitones. Range: -100 to 100. Positive values raise pitch, negative values lower it. |
| voice_modify_intensity | INT | 0-100–100 | Intensity/energy of the voice. Range: -100 to 100. Higher values create more energetic speech. |
| voice_modify_timbre | INT | 0-100–100 | Timbre adjustment. Range: -100 to 100. Affects the tonal quality of the voice. |
| max_cost_usdopt | FLOAT | 20.00–100000 | Safety cap. If the estimated cost of this run is above this many US dollars the node stops BEFORE sending anything to fal. 0 = no cap. |
| reuse_identical_runopt | BOOLEAN | true | If these exact inputs (the settings AND the same pictures/clips/audio) were already paid for and the files are still in output/fal, return them instead of paying again - also after a ComfyUI restart. Turn off for a fresh take with the same settings. |
| extra_jsonopt | STRING | Advanced: a JSON object merged into the request last, for any fal field this node has no box for. Example: {"seed": 7}. Leave empty normally. | |
| max_concurrentopt | INT | 11–32 | How many paid fal calls may run at the same time in one run, across every fal node. 1 (default) = one after another. This node starts its paid part only while fewer than this many fal calls are running. Once any fal node fails, the ones still waiting are not started, so nothing more is billed. |
Outputs (4)
| Name | Type | Description |
|---|---|---|
| audio | AUDIO | — |
| audio_path | STRING | — |
| duration_ms | INT | — |
| info | STRING | JSON: request id, cost estimate, saved files and fal's full answer. |