Nodes/MOSS-TTS 1.5/MOSS-TTS Voice Continue
ComfyUI Node

MOSS-TTS Voice Continue

Keep the same MOSS voice talking — without re-cloning

By eehrich·Created 3 months ago·Updated 3 days ago· 2
MOSS-TTS Voice Continue
  • moss_model
  • previous_audio
  • prev_tokens
  • audio
  • tokens_generated
  • full_audio
  • full_tokens
  • tokens
◄previous_text►
◄text►
◄languageEnglish►
◄audio_temperature1.70►
◄audio_top_p0.80►
◄audio_top_k25►
◄target_tokens0►
◄max_new_tokens4096►
◄seed42►
◄previous_tokens0►
◄head_trim_frames1►
◄target_overshoot_frames50►
◄audio_repetition_penalty1.00►
◄text_temperature1.00►
◄text_top_p1.00►
◄text_top_k50►
◄prefix_tail_trim_frames0►

MOSSVoiceContinue is the "don't make me re-clone it" node. You've generated a clip in some voice - via Speak, Voice Clone, or an earlier Continue - and you want the same voice to keep talking through the next paragraph. This node extends a prior MOSS clip with new text, inheriting the voice straight from the audio. No new reference clip, no re-clone, no seam. It's how you narrate a multi-paragraph passage without generating the whole thing at once.

The mechanism that bites people

MOSS is a prefix-continuation model. It doesn't just take audio and "say more" - it needs to know where it is in the script. So the node concatenates previous_text + " " + text into the full script, conditions on the prior audio, and generates audio for the new part. Which means previous_text must be the exact text that produced previous_audio - word-for-word, punctuation included. Get it wrong and you get Kauderwelsch, because the model can't align its position in the script. It's the single most common mistake, and the tooltip spells it out for a reason.

Inputs that matter

  • moss_model - from the loader.
  • previous_audio - the prior MOSS output.
  • previous_text - the exact text behind that audio. Non-negotiable.
  • text - the new lines, appended after previous_text.
  • previous_tokens - the exact frame count of the prefix. Wire the upstream node's tokens_generated output here for a precise handoff; leave at 0 and the node measures from the audio duration (off by at most one frame).
  • target_tokens - target length of the NEW segment only. The node adds the prefix internally because MOSS reads its hint as TOTAL (prefix + new) in continuation mode.
  • head_trim_frames - default 1 (~80 ms) trimmed from the start of the new audio. MOSS's conv codec lets the last prefix frame bleed into the start of the continuation; this removes it. Set to 0 if you hear a chopped syllable, higher if the bleed persists.

Outputs: audio (new segment only, for per-segment QC - you hear just the delta), tokens_generated (frames of the new segment), full_audio (prefix + new, concatenated and resampled to the model's native rate), and full_tokens (cumulative frame count). The chain pattern:

seg N:   Voice Continue -> audio, full_audio, full_tokens
seg N+1: previous_audio  <- seg N.full_audio
         previous_tokens <- seg N.full_tokens

The VRAM trap

Chaining segment N+1 with the whole cumulative history is the classic failure mode: VRAM drifts up about 1 GB per 20 s of history and eventually OOMs mid-scene. The fix the author documents: pass only the last segment's audio as previous_audio (and its full_tokens as previous_tokens) rather than the whole concatenation. Prefix-continuation semantics still hold - MOSS aligns the prefix at the end of previous_text inside the full script - just with a shorter history. Watch nvidia-smi during a long scene to spot the drift early.

Install & troubleshooting

Same pack as the rest - Manager, search "MOSS-TTS 1.5", or:

cd ComfyUI/custom_nodes
git clone https://github.com/eehrich/ComfyUI-MOSS-TTS-1.5.git MOSS-TTS-ComfyUI

Restart and load the model once first. A few notes from the README and source: an empty follow-up text is technically allowed but MOSS closes out almost immediately, so supply real lines. And like every generate node in the pack, a mild audio_repetition_penalty (1.05–1.15) is the cure if a segment starts droning or looping syllables. When target_tokens is set, target_overshoot_frames caps how far past it the model can run - a fuse against hangs on pathological input.

The neat trick: because the voice lives in the audio, not a separate reference, you can chain an arbitrary number of segments and each one inherits the voice for free. That's the entire point - one voice, one long take, no seams.

CategoryMOSS TTS 1.5

Inputs (20)

NameTypeDefaultDescription
moss_modelMOSS_MODELModel bundle produced by MOSS-TTS Load Model.
previous_textSTRINGThe exact text that produced 'previous_audio' / 'prev_tokens' -- its transcript. MOSS aligns the spoken reference against this text to find its 'where am I in the script?' state, so it must match word-for-word (punctuation matters). IF THE TEXT DOES NOT TRANSCRIBE THE AUDIO, THE OUTPUT IS GARBAGE -- not a degraded voice, gibberish. Trimmed the reference WAV? Trim this text to the same point. Pasted the wrong paragraph? Same result. The ONE deliberate exception is 'prefix_tail_trim_frames': there the audio ends earlier while this text stays whole, and MOSS simply says the last words again.
textSTRINGNew text to speak after the previous audio ends. Internally concatenated as: previous_text + ' ' + text -> full script. Must not be empty: with nothing to say MOSS never emits its end token and generates until max_new_tokens, so the node refuses it up front.
languageCOMBOEnglishLanguage hint for the follow-up text.
audio_temperatureFLOAT1.700.1–3Sampling temperature (default 1.7 = the 8B's generate() default).
audio_top_pFLOAT0.800–1Nucleus (top-p) sampling cutoff.
audio_top_kINT251–200Top-k sampling cutoff.
target_tokensINT00–65536Target length of the NEW continuation segment, in audio frames (12.5 fps). 0 = disabled (model decides via EOS). The node adds the measured prefix length internally, because MOSS reads its 'tokens' hint as TOTAL (prefix + new) in continuation mode. Chain a MOSS-TTS Estimate Tokens node on the FOLLOW-UP text to compute this.
max_new_tokensINT4096256–65536Safety cap on newly-generated audio frames. Same units as in Voice Clone: 12.5 fps -> 4096 caps continuation at ~5 min.
seedINT420–4294967295Random seed. Same seed + same inputs -> identical output.
previous_tokensINT00–65536Exact frame count of 'previous_audio'. Wire the 'tokens_generated' output of the preceding Speak / Voice Clone / Voice Continue node into this input for a precise handoff. Leave at 0 to measure from the audio duration (fine for externally-loaded WAVs, off by <=1 frame due to rounding). Ignored when 'prev_tokens' is wired: the code tensor already carries the exact length.
head_trim_framesINT10–10Extra frames to trim from the START of the new audio (1 frame = 80 ms at 12.5 fps). MOSS's decoder trims the prefix by SAMPLE proportion, and its conv-based codec has a receptive field that spans frame boundaries -- so the last prefix frame can bleed audibly into the start of the returned continuation. Default 1 (~80 ms) removes it in most cases. Set 0 to disable, higher if the bleed is longer.
target_overshoot_framesINT500–65536Runaway safety cap on the NEW continuation segment: when target_tokens > 0, MOSS may only exceed target_tokens by this many frames. Effective max_new_tokens = min(max_new_tokens, target_tokens + target_overshoot_frames). Default 50 = 4 s slack at 12.5 fps. Prevents long hangs on pathological inputs. Ignored when target_tokens = 0.
previous_audiooptAUDIOPrior MOSS output to continue from. Typically the AUDIO output of a preceding MOSS-TTS Voice Clone / Voice Continue node. Must be paired with the exact 'previous_text' that produced it -- a clip and a transcript that do not describe the same speech yield gibberish, so cut both to the same point or neither. Required unless 'prev_tokens' is wired -- it is only declared optional because ComfyUI has no way to express 'required unless that other input is connected'. Still worth wiring alongside prev_tokens if you want the 'full_audio' output: without it there is no prior waveform to prepend.
prev_tokensoptMOSS_TOKENSAudio codes of the previous segment (NOT the 'previous_tokens' frame COUNT below). When wired this REPLACES 'previous_audio': no codec re-encode, and the prefix length is taken exactly from the tensor. Feed the 'tokens' output of the preceding Voice Clone / Voice Continue node. RULE: 'previous_text' must transcribe EXACTLY what these codes contain -- shorten the token stream and you must shorten the transcript to the same point. A mismatched pair produces gibberish, not a slightly-off voice.
audio_repetition_penaltyoptFLOAT1.001–2Penalty on recently generated audio tokens (1.0 = off). Mild values (1.05-1.15) suppress the classic autoregressive TTS failure modes -- droning, tempo freeze, smeared or looping syllables -- while leaving normal prosody untouched (it only bites on pathological repeats). Values above ~1.3 can distort legitimately repeated sounds ('nein, nein, nein'). 8B: MOSS applies it to ALL codes in the sequence, the reference and the prefix included, pooled across codebooks -- with a long reference it works as a blanket push away from the cloned voice. Keep it at 1.0 there unless you hear a loop.
text_temperatureoptFLOAT1.000.1–3Temperature for the TEXT stream (alignment/pacing), NOT the acoustics. Default 1.0 = the 1.7B's own default (the 8B's generate() would use 1.5). Lower = steadier pacing/alignment; does not flatten the voice (that's audio_temperature).
text_top_poptFLOAT1.000–1Nucleus (top-p) cutoff for the TEXT stream. MOSS default 1.0 (off). Do not use 0.0: the 1.7B treats it as 'no filter', the 8B as near-greedy -- opposite results.
text_top_koptINT501–200Top-k cutoff for the TEXT stream. MOSS default 50.
prefix_tail_trim_framesoptINT0-1–256Cut this many frames off the END of the prefix before handing it to MOSS. 0 = off (unchanged behaviour), -1 = auto (n_vq-1 on the 8B, 0 on the 1.7B). ONLY relevant for the 8B build. Why: the 8B writes its codes in a DELAY PATTERN and the processor cuts the last n_vq-1 rows of the prefix so the model resumes mid-diagonal. The fine codes of the last 31 prefix frames are therefore missing and the model has to re-invent detail for audio that is already fixed -- the usual source of glitches in chained 8B continuation. Ending the prefix 31 frames earlier avoids that seam. Measured against the reference implementation: 31 is on par with it, 0 glitches, and 16/30/32/48 are all worse -- a one-frame-wide optimum at exactly n_vq-1. The price: previous_text still describes the FULL prefix, so MOSS re-speaks the trimmed ~2.5 s at the start of the new segment. That overlap is in 'audio' AND 'tokens', so 'full_audio' contains it twice -- cut it downstream before chaining (head_trim_frames stops at 10 frames, too short for this). On the 1.7B there is no delay pattern and no measurable gain: leave it at 0.

Outputs (5)

NameTypeDescription
audioAUDIONew segment only, head-trimmed. Use this for per-segment QC / preview -- you hear just the delta MOSS produced this call.
tokens_generatedINTFrames of the NEW segment only (frames, not samples). At 12.5 fps this equals duration_seconds * 12.5.
full_audioAUDIOCumulative audio: previous_audio + new segment concatenated at the model's rate, always as two channels (dual mono on the 8B). Wire this into the NEXT Continue's previous_audio when the same speaker keeps talking across segments. Falls back to the new segment alone when only prev_tokens (no previous_audio) is wired -- there is no prior waveform to prepend then.
full_tokensINTCumulative frame count: prefix + new. Wire into the next Continue's previous_tokens for a precise handoff without re-measurement. Equals 'tokens_generated' when no previous_audio is wired.
tokensMOSS_TOKENSAudio codes MOSS just generated, shape [frames, n_vq] at 12.5 fps (1 row = 80 ms). Feed into the next node's reference_tokens / prev_tokens (optionally through MOSS-TTS Concat Tokens) to keep the whole chain encode-free. These are the RAW emitted codes: they cover the untrimmed segment, so they include the frames removed by head_trim_frames from the AUDIO output.