Nodes/MOSS-TTS 1.5/MOSS-TTS Concat Tokens
ComfyUI Node

MOSS-TTS Concat Tokens

Joins two to four MOSS_TOKENS streams end to end. Built for sliding-window references: keep a fixed base anchor (the voice you cloned from) and append the tokens of the most recent segment(s), then feed the result into Voice Clone's 'reference_tokens'. That gives the same continuity as concatenating reference WAVs, without touching audio at all -- no decode, no re-encode, no resample. All inputs must come from the same model (identical n_vq).

By eehrich·Created 3 months ago·Updated about 9 hours ago· 2
MOSS-TTS Concat Tokens
  • tokens_a
  • tokens_b
  • tokens_c
  • tokens_d
  • tokens
  • frames
CategoryMOSS TTS 1.5

Inputs (4)

NameTypeDefaultDescription
tokens_aMOSS_TOKENSFirst stream. For a sliding window this is the base voice anchor.
tokens_bMOSS_TOKENSSecond stream, appended after tokens_a (e.g. the most recent segment).
tokens_coptMOSS_TOKENSOptional third stream, appended after tokens_b.
tokens_doptMOSS_TOKENSOptional fourth stream, appended after tokens_c.

Outputs (2)

NameTypeDescription
tokensMOSS_TOKENSConcatenated codes, shape [sum(frames), n_vq].
framesINTTotal frame count. Divide by 12.5 for seconds -- handy to keep a sliding-window reference inside a duration budget.