ComfyUI Node
MOSS-TTS Encode Tokens
Turns an AUDIO clip into MOSS_TOKENS (raw audio codes) with the model's own codec. Use this for the base voice reference so it never has to be encoded again: Voice Clone re-encodes its reference_audio on EVERY run, and with a long reference that encode can dominate the request -- it is a single codec pass that does not parallelise across concurrent requests, unlike generation itself. Encode once, wire the MOSS_TOKENS into reference_tokens (and/or save it with MOSS-TTS Save Tokens), and every later run starts straight at generation. Output is [frames, n_vq] at 12.5 fps, so 1 frame = 80 ms. Tokens are tied to the loaded model -- re-encode when you switch between the 1.7B and 8B build.
MOSS-TTS Encode Tokens
- moss_model
- audio
- tokens
- frames
CategoryMOSS TTS 1.5
Inputs (2)
| Name | Type | Default | Description |
|---|---|---|---|
| moss_model | MOSS_MODEL | Model bundle produced by MOSS-TTS Load Model. | |
| audio | AUDIO | Audio to encode. Resampled to the model's native rate and loudness-normalised by the processor exactly as the Voice Clone reference path does, so the resulting codes are interchangeable with it. Two rules the clip itself has to satisfy, both of which produce GIBBERISH (not a weaker voice) when broken: give it at least ~10 s -- around 5 s is not enough acoustic evidence for MOSS to lock onto the voice; and if these codes later serve as a Voice Continue prefix, whatever text you pass as 'previous_text' must transcribe THIS clip exactly -- trim the audio and you must trim the transcript to the same point. |
Outputs (2)
| Name | Type | Description |
|---|---|---|
| tokens | MOSS_TOKENS | Audio codes, shape [frames, n_vq]. Wire into Voice Clone 'reference_tokens', Voice Continue 'prev_tokens', MOSS-TTS Concat Tokens or MOSS-TTS Save Tokens. |
| frames | INT | Number of code frames. Divide by 12.5 for seconds. |