Nodes/ComfyUI-BytePlus-ModelArk/BytePlus Seed Speech ASR
ComfyUI Node

BytePlus Seed Speech ASR

Transcripts, timings and SRT without the Whisper dance

By byteplus-sa·Created 8 days ago·Updated about 8 hours ago· 3
BytePlus Seed Speech ASR
  • audio
  • context_image
  • text
  • utterances_json
  • srt
  • duration
◄modelseed-asr-fast►
◄audio_url►
◄languageauto►
◄enable_punctrue►
◄enable_itntrue►
◄enable_ddcfalse►
◄enable_speaker_infofalse►
◄hotwords►
◄context_text►
◄context_image_url►
◄enable_auto_langfalse►
◄enable_lidfalse►
◄enable_channel_splitfalse►
◄vad_segmentfalse►
◄end_window_size0►
◄output_zh_variantnone►
◄filter_system_sensitive_wordsfalse►
◄remove_words►
◄mask_words►
◄wrap_sensitive_wordsfalse►
◄audio_formatauto►

Transcription is the least glamorous node in this pack and the one that quietly enables the rest of it. The Podcast Clip template writes a two-voice script, synthesises it, then runs this node to produce SRT subtitles. The Multilingual Dubbing template transcribes a speech clip, translates it with the LLM, and re-voices it. Neither works without ASR, and if you've tried to get word-level timings out of a local Whisper node, you know why it's nice to have it as one call.

It's also the node where the fast/standard split matters most, because that split decides how your audio gets to the service.

How it works

model picks the path:

  • seed-asr-fast (default) sends a connected clip inline - one request, up to 2 hours or 100 MB. No upload, no login, done.
  • seed-asr-2.0 / seed-asr-1.0 submit and poll, handle up to 5 hours and more languages - but they only accept URLs. So a connected clip has to be uploaded to Comfy.org storage first, which means a Comfy.org login, or you supply a public URL.

Audio comes in via audio (a Load Audio node, for instance) or audio_url; one or the other. audio_format tells the service what container your URL is, and auto reads the extension - set it explicitly for URLs that don't have one.

The inputs worth setting

  • language defaults to auto, which covers Chinese, English and Chinese dialects. For anything else, name the language - and note the last sixteen languages on the list require the 2.x/1.0 models, not fast.
  • enable_punc (punctuation), enable_itn (numbers as digits) and enable_ddc (drop filler words and repetitions) are the readability trio; ddc is the one that turns "um, so, basically" into a usable transcript.
  • enable_speaker_info labels speakers - best with ten or fewer people, per the tooltip. enable_lid tags each utterance with its detected language, enable_channel_split treats stereo left and right as separate speakers, and enable_auto_lang detects the spoken language, overriding language.
  • hotwords is the accuracy lever for names, brands and jargon: one per line or comma-separated, up to 5000, and it's Chinese-English model only (auto or zh-CN).
  • context_text lets you give the model dialogue history or a scene description so it stops guessing at homophones, and context_image_url / context_image (ASR 2.0) gives it a picture for visual context - the pack resizes the connected image to fit the 500 KB limit before uploading it.
  • Segmentation: vad_segment splits on silence instead of meaning, end_window_size sets the silence in milliseconds that ends a sentence (300–5000, 0 = semantic), and output_zh_variant converts Chinese output to Traditional (traditional, tw, hk).
  • Cleanup: remove_words deletes terms, mask_words replaces them with *, filter_system_sensitive_words masks the built-in list, and wrap_sensitive_words wraps filtered words in backticks instead of hiding them - that last one is what you want if you need to see what was censored rather than silently lose it.

Outputs are text (the plain transcript), utterances_json (utterances with timings), srt (subtitle file contents, ready for a Save Text) and duration as a float.

Install and the key

cd ComfyUI/custom_nodes
git clone https://github.com/byteplus-sa/ComfyUI-BytePlus-ModelArk
pip install -r ComfyUI-BytePlus-ModelArk/requirements.txt

Restart (ComfyUI 0.31.0 or newer), or install from Manager by searching BytePlus ModelArk. Seed Speech has its own API key - same row in Settings → BytePlus, or BYTEPLUS_SEED_SPEECH_API_KEY in user/.env - and it runs in ap-southeast-1 only. ModelArk keys don't work here; you'll get Invalid X-Api-Key if you try.

Where people get burned

The login requirement on the standard models is the number one surprise. You connect Load Audio, run, and get told a Comfy.org login is needed - because seed-asr-2.0 and 1.0 only take URLs and your clip had to be uploaded somewhere first. seed-asr-fast avoids the whole problem; the tradeoff is the language list and the two-hour ceiling.

The other one is the misunderstanding about what ASR is for here. This isn't a captioning tool for describing images, and it isn't a diarisation service for a noisy twelve-person meeting - speaker labelling works best with ten or fewer clean voices, and heavy crosstalk is where it falls down like everything else. Feed it a voiceover, a podcast, a dubbed clip, and it does the job in one call, SRT included.

CategoryBytePlus ModelArk/Speech

Inputs (23)

NameTypeDefaultDescription
modelCOMBOseed-asr-fastseed-asr-fast: one request, audio up to 2 h / 100 MB. seed-asr-2.0 / 1.0: submit and poll, up to 5 h, more languages; connected audio is uploaded to Comfy.org storage.
audio_urlSTRINGPublic http(s) audio URL. Use this or the audio input (e.g. Load Audio).
languageCOMBOautoauto recognizes Chinese, English and Chinese dialects; pick a language for others. The last 16 languages need seed-asr-2.0 or 1.0.
enable_puncBOOLEANtrueAdd punctuation.
enable_itnBOOLEANtrueWrite numbers in digits (inverse text normalization).
enable_ddcBOOLEANfalseRemove filler words and repetitions.
enable_speaker_infoBOOLEANfalseLabel speakers (best with 10 or fewer).
hotwordsSTRINGNames or terms to favour, one per line or comma-separated (up to 5000). Chinese-English model only: language auto or zh-CN.
context_textSTRINGDialogue history or scene description, one entry per line, newest first (up to 20 entries / 800 tokens). Prefix a line with 'user:' or 'bot:' to mark the speaker. Chinese-English model only.
context_image_urlSTRINGASR 2.0 visual context: public JPEG/PNG URL (up to 500 KB). Or connect context_image.
enable_auto_langBOOLEANfalseDetect the spoken language automatically (overrides language).
enable_lidBOOLEANfalseseed-asr-2.0 / 1.0: add a detected-language label (lid_lang) to each utterance.
enable_channel_splitBOOLEANfalseRecognize the left and right channels of stereo audio separately (channel_id 1 / 2).
vad_segmentBOOLEANfalseSplit sentences on silence (VAD) instead of meaning.
end_window_sizeINT00–5000Silence in ms that ends a sentence (300-5000; 0 = semantic segmentation).
output_zh_variantCOMBOnoneConvert Chinese output to Traditional Chinese: traditional, tw (Taiwan) or hk (Hong Kong).
filter_system_sensitive_wordsBOOLEANfalseMask words from the built-in sensitive word list with *.
remove_wordsSTRINGWords to delete from the transcript, comma-separated.
mask_wordsSTRINGWords to replace with * in the transcript, comma-separated.
wrap_sensitive_wordsBOOLEANfalseWrap filtered words in backticks instead of hiding them silently.
audio_formatCOMBOautoContainer of audio_url. auto reads the file extension; the standard models read the file itself when the link has none. Set it if detection fails.
audiooptAUDIOAudio to transcribe, e.g. from Load Audio.
context_imageoptIMAGEASR 2.0 visual context image (uploaded to Comfy.org storage, resized to fit 500 KB).

Outputs (4)

NameTypeDescription
textSTRING—
utterances_jsonSTRING—
srtSTRING—
durationFLOAT—