Nodes/comfyui-arkennemasis/arkennemasis fal · MiniMax Speech 02 HD · $0.1/1k chars
ComfyUI Node

arkennemasis fal · MiniMax Speech 02 HD · $0.1/1k chars

Generate speech from text prompts and different voices using the MiniMax Speech-02 HD model, which leverages advanced AI techniques to create high-quality text-to-speech. Pricing (from fal, 2026-09-29): Your request will cost $0.1 per 1000 characters. Model page: https://fal.ai/models/fal-ai/minimax/speech-02-hd Key: FAL_KEY in the .env next to run_nvidia_gpu.bat. Results are saved in output/fal/fal-ai_minimax_speech-02-hd/.

By Hishamahmer·Created 2 months ago·Updated 4 days ago· 10
arkennemasis fal · MiniMax Speech 02 HD · $0.1/1k chars
    • audio
    • audio_path
    • duration_ms
    • info
    ◄text►
    ◄voice_setting_speed1.00►
    ◄voice_setting_vol1.00►
    ◄voice_setting_english_normalizationfalse►
    ◄voice_setting_pitch0►
    ◄voice_setting_emotion(not set)►
    ◄voice_setting_voice_idWise_Woman►
    ◄audio_setting_formatmp3►
    ◄audio_setting_bitrate128000►
    ◄audio_setting_sample_rate32000►
    ◄audio_setting_channel1►
    ◄language_boost(not set)►
    ◄output_formathex►
    ◄pronunciation_dict►
    ◄max_cost_usd20.0►
    ◄reuse_identical_runtrue►
    ◄extra_json►
    ◄max_concurrent1►
    Categoryarkennemasis/fal/Audio/MiniMax

    Inputs (18)

    NameTypeDefaultDescription
    textSTRINGText to convert to speech (max 5000 characters, minimum 1 non-whitespace character)
    voice_setting_speedFLOAT1.000.5–2Speech speed (0.5-2.0)
    voice_setting_volFLOAT1.000.01–10Volume (0-10)
    voice_setting_english_normalizationBOOLEANfalseEnables English text normalization to improve number reading performance, with a slight increase in latency
    voice_setting_pitchINT0-12–12Voice pitch (-12 to 12)
    voice_setting_emotionCOMBO(not set)Emotion of the generated speech '(not set)' leaves it to the model.
    voice_setting_voice_idSTRINGWise_WomanPredefined voice ID to use for synthesis Empty leaves it out.
    audio_setting_formatCOMBOmp3Audio format
    audio_setting_bitrateCOMBO128000Bitrate of generated audio
    audio_setting_sample_rateCOMBO32000Sample rate of generated audio
    audio_setting_channelCOMBO1Number of audio channels (1=mono, 2=stereo)
    language_boostCOMBO(not set)Enhance recognition of specified languages and dialects '(not set)' leaves it to the model.
    output_formatCOMBOhexFormat of the output content (non-streaming only)
    pronunciation_dictSTRINGCustom pronunciation dictionary for text replacement JSON. Empty leaves it out.
    max_cost_usdoptFLOAT20.00–100000Safety cap. If the estimated cost of this run is above this many US dollars the node stops BEFORE sending anything to fal. 0 = no cap.
    reuse_identical_runoptBOOLEANtrueIf these exact inputs (the settings AND the same pictures/clips/audio) were already paid for and the files are still in output/fal, return them instead of paying again - also after a ComfyUI restart. Turn off for a fresh take with the same settings.
    extra_jsonoptSTRINGAdvanced: a JSON object merged into the request last, for any fal field this node has no box for. Example: {"seed": 7}. Leave empty normally.
    max_concurrentoptINT11–32How many paid fal calls may run at the same time in one run, across every fal node. 1 (default) = one after another. This node starts its paid part only while fewer than this many fal calls are running. Once any fal node fails, the ones still waiting are not started, so nothing more is billed.

    Outputs (4)

    NameTypeDescription
    audioAUDIO—
    audio_pathSTRING—
    duration_msINT—
    infoSTRINGJSON: request id, cost estimate, saved files and fal's full answer.