Nodes/MOSS-TTS 1.5/MOSS-TTS Decode Tokens
ComfyUI Node

MOSS-TTS Decode Tokens

Inverse of MOSS-TTS Encode Tokens: sends MOSS audio codes back through the model's own vocoder and returns a normal ComfyUI AUDIO dict. This completes the token API (encode / decode / concat / save / load) and is what makes a token stream auditable -- listen to what a saved token file actually holds, hear the exact reference window a MOSS-TTS Concat Tokens result builds, or check a 'tokens' output frame for frame, all WITHOUT generating anything. Especially useful before a long run: a reference that sounds wrong here will clone wrong. Decoding is a pure codec pass, no sampling -- deterministic, no seed. Codes are model-specific, so decode with the model that produced them (n_vq is validated).

By eehrich·Created 3 months ago·Updated about 9 hours ago· 2
MOSS-TTS Decode Tokens
  • moss_model
  • tokens
  • audio
  • frames
◄return_stereotrue►
CategoryMOSS TTS 1.5

Inputs (3)

NameTypeDefaultDescription
moss_modelMOSS_MODELModel bundle produced by MOSS-TTS Load Model. The vocoder is part of the model, so decode with the same variant that encoded/emitted the codes -- it also defines the output sample rate (48 kHz for the 1.7B Local-Transformer, 24 kHz for the 8B).
tokensMOSS_TOKENSCodes to turn back into audio, shape [frames, n_vq] at 12.5 fps (1 row = 80 ms). Any MOSS_TOKENS source works: MOSS-TTS Encode / Concat / Load Tokens, or the 'tokens' output of Speak / Voice Clone / Voice Continue.
return_stereooptBOOLEANtrueChannel layout, passed straight to the processor's decode_audio_codes(return_stereo=...). True (default) keeps the codec's native stereo -- identical to what every generate node in this pack outputs, so a decoded 'tokens' output lines up with its own 'audio'. False averages the codec channels into one mono channel: smaller, but no longer bit-identical to the generate nodes' audio. 1.7B only: the 8B codec is mono and ignores this switch.

Outputs (2)

NameTypeDescription
audioAUDIODecoded audio at the model's native sample rate -- stereo on the 1.7B unless return_stereo is off, always mono on the 8B -- ready for PreviewAudio / SaveAudio.
framesINTNumber of code frames decoded. Divide by 12.5 for seconds.