BytePlus Seed Speech ASR
Transcripts, timings and SRT without the Whisper dance
- audio
- context_image
- text
- utterances_json
- srt
- duration
Transcription is the least glamorous node in this pack and the one that quietly enables the rest of it. The Podcast Clip template writes a two-voice script, synthesises it, then runs this node to produce SRT subtitles. The Multilingual Dubbing template transcribes a speech clip, translates it with the LLM, and re-voices it. Neither works without ASR, and if you've tried to get word-level timings out of a local Whisper node, you know why it's nice to have it as one call.
It's also the node where the fast/standard split matters most, because that split decides how your audio gets to the service.
How it works
model picks the path:
seed-asr-fast(default) sends a connected clip inline - one request, up to 2 hours or 100 MB. No upload, no login, done.seed-asr-2.0/seed-asr-1.0submit and poll, handle up to 5 hours and more languages - but they only accept URLs. So a connected clip has to be uploaded to Comfy.org storage first, which means a Comfy.org login, or you supply a public URL.
Audio comes in via audio (a Load Audio node, for instance) or audio_url; one or the other. audio_format tells the service what container your URL is, and auto reads the extension - set it explicitly for URLs that don't have one.
The inputs worth setting
languagedefaults toauto, which covers Chinese, English and Chinese dialects. For anything else, name the language - and note the last sixteen languages on the list require the 2.x/1.0 models, not fast.enable_punc(punctuation),enable_itn(numbers as digits) andenable_ddc(drop filler words and repetitions) are the readability trio;ddcis the one that turns "um, so, basically" into a usable transcript.enable_speaker_infolabels speakers - best with ten or fewer people, per the tooltip.enable_lidtags each utterance with its detected language,enable_channel_splittreats stereo left and right as separate speakers, andenable_auto_langdetects the spoken language, overridinglanguage.hotwordsis the accuracy lever for names, brands and jargon: one per line or comma-separated, up to 5000, and it's Chinese-English model only (autoorzh-CN).context_textlets you give the model dialogue history or a scene description so it stops guessing at homophones, andcontext_image_url/context_image(ASR 2.0) gives it a picture for visual context - the pack resizes the connected image to fit the 500 KB limit before uploading it.- Segmentation:
vad_segmentsplits on silence instead of meaning,end_window_sizesets the silence in milliseconds that ends a sentence (300–5000,0= semantic), andoutput_zh_variantconverts Chinese output to Traditional (traditional,tw,hk). - Cleanup:
remove_wordsdeletes terms,mask_wordsreplaces them with*,filter_system_sensitive_wordsmasks the built-in list, andwrap_sensitive_wordswraps filtered words in backticks instead of hiding them - that last one is what you want if you need to see what was censored rather than silently lose it.
Outputs are text (the plain transcript), utterances_json (utterances with timings), srt (subtitle file contents, ready for a Save Text) and duration as a float.
Install and the key
cd ComfyUI/custom_nodes
git clone https://github.com/byteplus-sa/ComfyUI-BytePlus-ModelArk
pip install -r ComfyUI-BytePlus-ModelArk/requirements.txt
Restart (ComfyUI 0.31.0 or newer), or install from Manager by searching BytePlus ModelArk. Seed Speech has its own API key - same row in Settings → BytePlus, or BYTEPLUS_SEED_SPEECH_API_KEY in user/.env - and it runs in ap-southeast-1 only. ModelArk keys don't work here; you'll get Invalid X-Api-Key if you try.
Where people get burned
The login requirement on the standard models is the number one surprise. You connect Load Audio, run, and get told a Comfy.org login is needed - because seed-asr-2.0 and 1.0 only take URLs and your clip had to be uploaded somewhere first. seed-asr-fast avoids the whole problem; the tradeoff is the language list and the two-hour ceiling.
The other one is the misunderstanding about what ASR is for here. This isn't a captioning tool for describing images, and it isn't a diarisation service for a noisy twelve-person meeting - speaker labelling works best with ten or fewer clean voices, and heavy crosstalk is where it falls down like everything else. Feed it a voiceover, a podcast, a dubbed clip, and it does the job in one call, SRT included.
Inputs (23)
| Name | Type | Default | Description |
|---|---|---|---|
| model | COMBO | seed-asr-fast | seed-asr-fast: one request, audio up to 2 h / 100 MB. seed-asr-2.0 / 1.0: submit and poll, up to 5 h, more languages; connected audio is uploaded to Comfy.org storage. |
| audio_url | STRING | Public http(s) audio URL. Use this or the audio input (e.g. Load Audio). | |
| language | COMBO | auto | auto recognizes Chinese, English and Chinese dialects; pick a language for others. The last 16 languages need seed-asr-2.0 or 1.0. |
| enable_punc | BOOLEAN | true | Add punctuation. |
| enable_itn | BOOLEAN | true | Write numbers in digits (inverse text normalization). |
| enable_ddc | BOOLEAN | false | Remove filler words and repetitions. |
| enable_speaker_info | BOOLEAN | false | Label speakers (best with 10 or fewer). |
| hotwords | STRING | Names or terms to favour, one per line or comma-separated (up to 5000). Chinese-English model only: language auto or zh-CN. | |
| context_text | STRING | Dialogue history or scene description, one entry per line, newest first (up to 20 entries / 800 tokens). Prefix a line with 'user:' or 'bot:' to mark the speaker. Chinese-English model only. | |
| context_image_url | STRING | ASR 2.0 visual context: public JPEG/PNG URL (up to 500 KB). Or connect context_image. | |
| enable_auto_lang | BOOLEAN | false | Detect the spoken language automatically (overrides language). |
| enable_lid | BOOLEAN | false | seed-asr-2.0 / 1.0: add a detected-language label (lid_lang) to each utterance. |
| enable_channel_split | BOOLEAN | false | Recognize the left and right channels of stereo audio separately (channel_id 1 / 2). |
| vad_segment | BOOLEAN | false | Split sentences on silence (VAD) instead of meaning. |
| end_window_size | INT | 00–5000 | Silence in ms that ends a sentence (300-5000; 0 = semantic segmentation). |
| output_zh_variant | COMBO | none | Convert Chinese output to Traditional Chinese: traditional, tw (Taiwan) or hk (Hong Kong). |
| filter_system_sensitive_words | BOOLEAN | false | Mask words from the built-in sensitive word list with *. |
| remove_words | STRING | Words to delete from the transcript, comma-separated. | |
| mask_words | STRING | Words to replace with * in the transcript, comma-separated. | |
| wrap_sensitive_words | BOOLEAN | false | Wrap filtered words in backticks instead of hiding them silently. |
| audio_format | COMBO | auto | Container of audio_url. auto reads the file extension; the standard models read the file itself when the link has none. Set it if detection fails. |
| audioopt | AUDIO | Audio to transcribe, e.g. from Load Audio. | |
| context_imageopt | IMAGE | ASR 2.0 visual context image (uploaded to Comfy.org storage, resized to fit 500 KB). |
Outputs (4)
| Name | Type | Description |
|---|---|---|
| text | STRING | — |
| utterances_json | STRING | — |
| srt | STRING | — |
| duration | FLOAT | — |