ComfyUI Node
arkennemasis fal · MiniMax Speech 02 Turbo · $0.06/1k chars
Generate fast speech from text prompts and different voices using the MiniMax Speech-02 Turbo model, which leverages advanced AI techniques to create high-quality text-to-speech. Pricing (from fal, 2026-09-29): Your request will cost $0.06 per 1000 character. Model page: https://fal.ai/models/fal-ai/minimax/speech-02-turbo Key: FAL_KEY in the .env next to run_nvidia_gpu.bat. Results are saved in output/fal/fal-ai_minimax_speech-02-turbo/.
arkennemasis fal · MiniMax Speech 02 Turbo · $0.06/1k chars
- audio
- audio_path
- duration_ms
- info
◄text►
◄voice_setting_speed1.00►
◄voice_setting_vol1.00►
◄voice_setting_english_normalizationfalse►
◄voice_setting_pitch0►
◄voice_setting_emotion(not set)►
◄voice_setting_voice_idWise_Woman►
◄audio_setting_formatmp3►
◄audio_setting_bitrate128000►
◄audio_setting_sample_rate32000►
◄audio_setting_channel1►
◄language_boost(not set)►
◄output_formathex►
◄pronunciation_dict►
◄max_cost_usd20.0►
◄reuse_identical_runtrue►
◄extra_json►
◄max_concurrent1►
Categoryarkennemasis/fal/Audio/MiniMax
Inputs (18)
| Name | Type | Default | Description |
|---|---|---|---|
| text | STRING | Text to convert to speech (max 5000 characters, minimum 1 non-whitespace character) | |
| voice_setting_speed | FLOAT | 1.000.5–2 | Speech speed (0.5-2.0) |
| voice_setting_vol | FLOAT | 1.000.01–10 | Volume (0-10) |
| voice_setting_english_normalization | BOOLEAN | false | Enables English text normalization to improve number reading performance, with a slight increase in latency |
| voice_setting_pitch | INT | 0-12–12 | Voice pitch (-12 to 12) |
| voice_setting_emotion | COMBO | (not set) | Emotion of the generated speech '(not set)' leaves it to the model. |
| voice_setting_voice_id | STRING | Wise_Woman | Predefined voice ID to use for synthesis Empty leaves it out. |
| audio_setting_format | COMBO | mp3 | Audio format |
| audio_setting_bitrate | COMBO | 128000 | Bitrate of generated audio |
| audio_setting_sample_rate | COMBO | 32000 | Sample rate of generated audio |
| audio_setting_channel | COMBO | 1 | Number of audio channels (1=mono, 2=stereo) |
| language_boost | COMBO | (not set) | Enhance recognition of specified languages and dialects '(not set)' leaves it to the model. |
| output_format | COMBO | hex | Format of the output content (non-streaming only) |
| pronunciation_dict | STRING | Custom pronunciation dictionary for text replacement JSON. Empty leaves it out. | |
| max_cost_usdopt | FLOAT | 20.00–100000 | Safety cap. If the estimated cost of this run is above this many US dollars the node stops BEFORE sending anything to fal. 0 = no cap. |
| reuse_identical_runopt | BOOLEAN | true | If these exact inputs (the settings AND the same pictures/clips/audio) were already paid for and the files are still in output/fal, return them instead of paying again - also after a ComfyUI restart. Turn off for a fresh take with the same settings. |
| extra_jsonopt | STRING | Advanced: a JSON object merged into the request last, for any fal field this node has no box for. Example: {"seed": 7}. Leave empty normally. | |
| max_concurrentopt | INT | 11–32 | How many paid fal calls may run at the same time in one run, across every fal node. 1 (default) = one after another. This node starts its paid part only while fewer than this many fal calls are running. Once any fal node fails, the ones still waiting are not started, so nothing more is billed. |
Outputs (4)
| Name | Type | Description |
|---|---|---|
| audio | AUDIO | — |
| audio_path | STRING | — |
| duration_ms | INT | — |
| info | STRING | JSON: request id, cost estimate, saved files and fal's full answer. |