ComfyUI Node
MOSS-TTS Decode Tokens
Inverse of MOSS-TTS Encode Tokens: sends MOSS audio codes back through the model's own vocoder and returns a normal ComfyUI AUDIO dict. This completes the token API (encode / decode / concat / save / load) and is what makes a token stream auditable -- listen to what a saved token file actually holds, hear the exact reference window a MOSS-TTS Concat Tokens result builds, or check a 'tokens' output frame for frame, all WITHOUT generating anything. Especially useful before a long run: a reference that sounds wrong here will clone wrong. Decoding is a pure codec pass, no sampling -- deterministic, no seed. Codes are model-specific, so decode with the model that produced them (n_vq is validated).
MOSS-TTS Decode Tokens
- moss_model
- tokens
- audio
- frames
◄return_stereotrue►
CategoryMOSS TTS 1.5
Inputs (3)
| Name | Type | Default | Description |
|---|---|---|---|
| moss_model | MOSS_MODEL | Model bundle produced by MOSS-TTS Load Model. The vocoder is part of the model, so decode with the same variant that encoded/emitted the codes -- it also defines the output sample rate (48 kHz for the 1.7B Local-Transformer, 24 kHz for the 8B). | |
| tokens | MOSS_TOKENS | Codes to turn back into audio, shape [frames, n_vq] at 12.5 fps (1 row = 80 ms). Any MOSS_TOKENS source works: MOSS-TTS Encode / Concat / Load Tokens, or the 'tokens' output of Speak / Voice Clone / Voice Continue. | |
| return_stereoopt | BOOLEAN | true | Channel layout, passed straight to the processor's decode_audio_codes(return_stereo=...). True (default) keeps the codec's native stereo -- identical to what every generate node in this pack outputs, so a decoded 'tokens' output lines up with its own 'audio'. False averages the codec channels into one mono channel: smaller, but no longer bit-identical to the generate nodes' audio. 1.7B only: the 8B codec is mono and ignores this switch. |
Outputs (2)
| Name | Type | Description |
|---|---|---|
| audio | AUDIO | Decoded audio at the model's native sample rate -- stereo on the 1.7B unless return_stereo is off, always mono on the 8B -- ready for PreviewAudio / SaveAudio. |
| frames | INT | Number of code frames decoded. Divide by 12.5 for seconds. |