Nodes/Tikpan Official Nodes/音频|speech-2.8-hd 高清语音合成
ComfyUI Node

音频|speech-2.8-hd 高清语音合成

Studio-grade TTS with 61 voices, pauses, and sound effects

By htrert·Created 5 months ago·Updated 2 months ago· 1
音频|speech-2.8-hd 高清语音合成
    • 📁_音频路径
    • 🔗_音频链接
    • 🆔_任务ID
    • 🆔_文件ID
    • 💰_计费字符数
    • 📋_状态日志
    • 🎧_音频流
    💵_福利_💵
    获取密钥请访问
    API_密钥sk-
    合成文本欢迎使用 Tikpan speech-2.8-hd。你可以在文本中加入 <#0.5#> 控制停顿,也可以使用 (laughs) 或 (sighs) 这类音效标签。
    模型speech-2.8-hd
    调用方式同步语音 /t2a_v2
    音色中文普通话|可靠管理者|稳重商务男声,适合企业旁白|Chinese (Mandarin)_Reliable_Executive
    语言增强auto
    语速1.00
    音量1.0
    音调0
    情绪默认不传
    采样率32000
    比特率128000
    音频格式mp3
    声道数1
    POST重试策略幂等键轻重试
    校验HTTPS证书true
    自定义voice_id
    同步返回格式hex
    发音字典_tone_每行一条
    音色混合_JSON
    声音效果
    音色修饰_pitch0
    音色修饰_intensity0
    音色修饰_timbre0
    开启字幕false
    字幕类型sentence
    最长等待秒数900
    轮询间隔秒数5
    复用本地缓存true
    跳过错误false

    MiniMax's speech-2.8-hd is their high-definition text-to-speech, and this node is a remarkably complete wrapper around it - 61 voices, per-language pronunciation boosting, speed/pitch/volume, emotion tags, a tone dictionary for name pronunciation, even a voice-blending JSON input. If you've ever fought a TTS node that gives you one model, one voice, and one rate, this is the opposite problem: the options are the feature. It's aimed at the pack's core audience of Chinese e-commerce content producers, but the voice list and language controls make it work for English narration too.

    The mechanism under the hood is MiniMax's T2A service reached through Tikpan's relay (/minimax/v1/t2a_v2 synchronous, or the async variant for long text), and the node does a couple of smart things for you. The default return mode is hex - the audio comes back inline as hex and gets decoded locally into a proper audio file, which is more reliable than chasing a hosted URL. And it defaults 复用本地缓存 to on: identical text plus identical params hit a local cache and return the saved audio instead of re-billing you. That last one is genuinely considerate - failed or repeated runs are how people accidentally burn TTS credit.

    The inputs that matter (and the syntax worth knowing)

    • 合成文本 - the text, and it supports MiniMax's control syntax: <#0.5#> inserts a 0.5-second pause, and (laughs), (sighs), (coughs) add sound effects. This is the single best lever for making TTS sound like a person instead of a robot.
    • 音色 - 61 voices, each labeled with gender/age/style in Chinese. The default is a Mandarin executive male voice.
    • 语速 / 音量 / 音调 - rate and level multipliers, plus semitone pitch.
    • 情绪 - happy/sad/angry etc.; leave at default to skip.
    • 语言增强 - boost a specific language's pronunciation, auto otherwise.
    • 采样率 / 比特率 / 音频格式 / 声道数 - output quality knobs; mp3/32000/128000 mono is a fine starting point.
    • 开启字幕 - returns a subtitle timeline alongside the audio if you want to burn captions.
    • 自定义voice_id, 音色混合_JSON, 发音字典, 声音效果 - the advanced lane: cloned/designed voices, voice blending with weights, per-word tone dict, and effects like spacious_echo.

    Outputs are 🎧_音频流 (AUDIO), 📁_音频路径, 🔗_音频链接, 🆔_任务ID, 🆔_文件ID, 💰_计费字符数, and 📋_状态日志 - wire the AUDIO stream into a save/preview node and keep 计费字符数 handy for cost tracking.

    Install and first run

    cd ComfyUI/custom_nodes
    git clone https://github.com/htrert/ComfyUI-Tikpan-Pro
    

    Restart (or Manager → "Tikpan"), paste your sk- key, type a sentence with a <#0.5#> pause and a (laughs), run. Nothing downloads; the model is server-side. Then start exploring voices - that's where most of the fun is.

    Common issues

    The pack-wide 401/402/429 family applies. TTS-specific: check 计费字符数 before long re-runs (repeated long text is how you discover the pricing), and keep 复用本地缓存 on so identical re-runs don't re-bill. If a sync call times out on long text, switch 调用方式 to async and collect via the task ID. If 🎧_音频流 won't preview but 📁_音频路径 has a file, the render succeeded - the preview gap is local. And the 校验HTTPS证书 toggle defaults to on here, so local-proxy setups that trip SSL should look there first.

    Category👑 Tikpan 官方独家节点/03 音频 Audio

    Inputs (32)

    NameTypeDefaultDescription
    💵_福利_💵COMBO1 options: 🔥 speech-2.8-hd 高清语音合成 | 自定义音色/复刻音色按上游规则额外计费
    获取密钥请访问COMBO1 options: 👉 https://tikpan.com (官方授权 Key 获取入口)
    API_密钥STRINGsk-Tikpan 平台的 API 密钥,以 sk- 开头,从 https://tikpan.com 获取
    合成文本STRING欢迎使用 Tikpan speech-2.8-hd。你可以在文本中加入 <#0.5#> 控制停顿,也可以使用 (laughs) 或 (sighs) 这类音效标签。需要被合成为语音的文本;支持 <#秒数#> 停顿、(laughs) 音效标签
    模型COMBOspeech-2.8-hd本节点使用的语音合成模型
    调用方式COMBO同步语音 /t2a_v2同步=直接等待返回;异步=适合长文本,配合任务查询节点
    音色COMBO中文普通话|可靠管理者|稳重商务男声,适合企业旁白|Chinese (Mandarin)_Reliable_Executive选择说话人音色:每个音色对应不同的性别/年龄/风格
    语言增强COMBOauto强化某种语言的发音;auto=自动识别
    语速FLOAT1.000.5–2语速倍率:1.0 标准,>1 更快,<1 更慢
    音量FLOAT1.00–10音量倍率:1.0 标准;过高可能爆音
    音调INT0-12–12调高半音:正数变尖,负数变沉
    情绪COMBO默认不传情绪标签:happy 欢快 / sad 悲伤 / angry 愤怒 等
    采样率COMBO32000输出音频采样率(Hz);越高越清晰但文件越大
    比特率COMBO128000音频比特率;越高音质越好文件越大
    音频格式COMBOmp3音频编码:mp3 通用,wav 无损未压缩,flac 无损压缩
    声道数COMBO11=单声道(更小),2=立体声
    POST重试策略COMBO幂等键轻重试网络异常重试方式;带幂等键更安全
    校验HTTPS证书BOOLEANtrue默认开启;遇到本地证书问题再关闭(不推荐关闭)
    自定义voice_idoptSTRING选择“自定义 voice_id”时填写;也兼容复刻音色、音色设计和官方新上线 voice_id。
    同步返回格式optCOMBOhex同步模式返回音频的方式:hex 内联十六进制;url 云端链接
    发音字典_tone_每行一条optSTRING示例:燕少飞/(yan4)(shao3)(fei1) 或 omg/oh my god
    音色混合_JSONoptSTRING高级参数,示例:[{"voice_id":"xxx","weight":70},{"voice_id":"yyy","weight":30}]
    声音效果optSTRING示例:spacious_echo。留空则不传 sound_effects。
    音色修饰_pitchoptINT0-100–100音高微调:正数变尖,负数变沉
    音色修饰_intensityoptINT0-100–100情感强度微调
    音色修饰_timbreoptINT0-100–100音色质感微调
    开启字幕optBOOLEANfalse开启后会同时返回字幕时间轴
    字幕类型optCOMBOsentence字幕粒度:sentence 按句;word 按词
    最长等待秒数optINT90030–7200异步任务等待上限秒数
    轮询间隔秒数optINT53–60异步任务轮询间隔秒数
    复用本地缓存optBOOLEANtrue相同文本和参数命中缓存时直接返回本地音频,减少误重复扣费。
    跳过错误optBOOLEANfalse批量工作流可开启。失败时返回空音频和错误日志,不中断整个工作流。

    Outputs (7)

    NameTypeDescription
    📁_音频路径STRING
    🔗_音频链接STRING
    🆔_任务IDSTRING
    🆔_文件IDSTRING
    💰_计费字符数STRING
    📋_状态日志STRING
    🎧_音频流AUDIO