Nodes/comfyui-timesaver/TS RT Prompt Enhancer
ComfyUI Node

TS RT Prompt Enhancer

Gemma 4 on LiteRT-LM as a graph node: enhance prompts, describe pictures and sound, write out speech, translate — driven by the same presets as TS Super Prompt RT.

By AlexYez·Created 2 years ago·Updated about 13 hours ago· 15
TS RT Prompt Enhancer
  • images
  • audio
  • text
◄modelGemma 4 E4B (3.4 GB)►
◄system_presetPrompts enhance►
◄prompt►
◄seed0►
◄max_new_tokens0►
◄keep_loadedfalse►
◄enabletrue►
◄audio_modelisten►
◄custom_system_prompt—►
CategoryTS/LLM

Inputs (11)

NameTypeDefaultDescription
modelCOMBOGemma 4 E4B (3.4 GB)E4B writes better, E2B is about twice as fast and lighter. Downloaded on first use into models/LLM/litert.
system_presetCOMBOPrompts enhanceWhat the model is asked to do. The same presets as TS Super Prompt RT and TS Qwen 3. 'Your instruction' uses the connected custom_system_prompt.
promptSTRINGYour idea, text or question. May be left empty when pictures or a recording are connected.
seedINT00–18446744073709550000Same seed, same settings, same answer. Change it for a different wording.
max_new_tokensINT00–4096Longest answer in tokens. 0 = the preset's own limit. The model stops there even mid-sentence — it is a guard against a runaway answer.
keep_loadedBOOLEANfalseKeep Gemma in memory between runs (faster repeats). ComfyUI CANNOT see this memory — it runs on WebGPU, not CUDA — and will plan other models as if it were free. Leave off unless the card has room to spare.
enableBOOLEANtrueOff: the prompt passes through unchanged and no model is loaded.
audio_modeCOMBOlistenWhat to do with a connected recording. listen — the model hears it itself: music, sounds, a short phrase. Only the first 30 s: that is the longest clip the model is trained on. transcribe — speech of any length (up to 5 min) is written out in 30 s pieces and added to the prompt. Tuned for Russian speech with English terms.
imagesoptIMAGEOne picture, two, or a whole clip's frames. At most 4 go in: a longer batch is sampled evenly from first to last. Each is shrunk to 1024 px — every pixel costs context.
audiooptAUDIOOptional recording — a song, a sound, a voice note, a clip's soundtrack. audio_mode decides whether the model listens to it or transcribes it.
custom_system_promptoptSTRINGYour own system prompt. Used when system_preset is 'Your instruction'.

Outputs (1)

NameTypeDescription
textSTRINGThe model's answer.