ComfyUI Node
MOSS-TTS Concat Tokens
Joins two to four MOSS_TOKENS streams end to end. Built for sliding-window references: keep a fixed base anchor (the voice you cloned from) and append the tokens of the most recent segment(s), then feed the result into Voice Clone's 'reference_tokens'. That gives the same continuity as concatenating reference WAVs, without touching audio at all -- no decode, no re-encode, no resample. All inputs must come from the same model (identical n_vq).
MOSS-TTS Concat Tokens
- tokens_a
- tokens_b
- tokens_c
- tokens_d
- tokens
- frames
CategoryMOSS TTS 1.5
Inputs (4)
| Name | Type | Default | Description |
|---|---|---|---|
| tokens_a | MOSS_TOKENS | First stream. For a sliding window this is the base voice anchor. | |
| tokens_b | MOSS_TOKENS | Second stream, appended after tokens_a (e.g. the most recent segment). | |
| tokens_copt | MOSS_TOKENS | Optional third stream, appended after tokens_b. | |
| tokens_dopt | MOSS_TOKENS | Optional fourth stream, appended after tokens_c. |
Outputs (2)
| Name | Type | Description |
|---|---|---|
| tokens | MOSS_TOKENS | Concatenated codes, shape [sum(frames), n_vq]. |
| frames | INT | Total frame count. Divide by 12.5 for seconds -- handy to keep a sliding-window reference inside a duration budget. |