ComfyUI Node
TS RT Prompt Enhancer
Gemma 4 on LiteRT-LM as a graph node: enhance prompts, describe pictures and sound, write out speech, translate — driven by the same presets as TS Super Prompt RT.
TS RT Prompt Enhancer
- images
- audio
- text
◄modelGemma 4 E4B (3.4 GB)►
◄system_presetPrompts enhance►
◄prompt►
◄seed0►
◄max_new_tokens0►
◄keep_loadedfalse►
◄enabletrue►
◄audio_modelisten►
◄custom_system_prompt—►
CategoryTS/LLM
Inputs (11)
| Name | Type | Default | Description |
|---|---|---|---|
| model | COMBO | Gemma 4 E4B (3.4 GB) | E4B writes better, E2B is about twice as fast and lighter. Downloaded on first use into models/LLM/litert. |
| system_preset | COMBO | Prompts enhance | What the model is asked to do. The same presets as TS Super Prompt RT and TS Qwen 3. 'Your instruction' uses the connected custom_system_prompt. |
| prompt | STRING | Your idea, text or question. May be left empty when pictures or a recording are connected. | |
| seed | INT | 00–18446744073709550000 | Same seed, same settings, same answer. Change it for a different wording. |
| max_new_tokens | INT | 00–4096 | Longest answer in tokens. 0 = the preset's own limit. The model stops there even mid-sentence — it is a guard against a runaway answer. |
| keep_loaded | BOOLEAN | false | Keep Gemma in memory between runs (faster repeats). ComfyUI CANNOT see this memory — it runs on WebGPU, not CUDA — and will plan other models as if it were free. Leave off unless the card has room to spare. |
| enable | BOOLEAN | true | Off: the prompt passes through unchanged and no model is loaded. |
| audio_mode | COMBO | listen | What to do with a connected recording. listen — the model hears it itself: music, sounds, a short phrase. Only the first 30 s: that is the longest clip the model is trained on. transcribe — speech of any length (up to 5 min) is written out in 30 s pieces and added to the prompt. Tuned for Russian speech with English terms. |
| imagesopt | IMAGE | One picture, two, or a whole clip's frames. At most 4 go in: a longer batch is sampled evenly from first to last. Each is shrunk to 1024 px — every pixel costs context. | |
| audioopt | AUDIO | Optional recording — a song, a sound, a voice note, a clip's soundtrack. audio_mode decides whether the model listens to it or transcribes it. | |
| custom_system_promptopt | STRING | Your own system prompt. Used when system_preset is 'Your instruction'. |
Outputs (1)
| Name | Type | Description |
|---|---|---|
| text | STRING | The model's answer. |