Nodes/was-node-suite-comfyui/Fast Generate Text
ComfyUI Node Runs on cloud

Fast Generate Text

Your Text Encoder Is Already an LLM. Use It.

By WASasquatch·Created 4 years ago·Updated a day ago· 1,864
Fast Generate Text
  • clip
  • image
  • video
  • audio
  • generated_text
◄prompt►
◄max_length512►
◄sampling_mode▾►
◄thinkingfalse►
◄use_default_templatetrue►
◄mtpauto►
◄fast_decodetrue►

In 2026 your text encoder is a language model. Z-Image and Flux 2 Klein 4B encode with Qwen3-4B, LTX-2 uses Gemma 3 12B, and those files already sit in models/text_encoders doing nothing but turning prompts into conditioning. ComfyUI's core Generate Text node lets you talk to one of them directly. Fast Generate Text does the same job on ComfyUI's fast decode path - several times quicker on exactly the models core leaves on the slow one.

The community's read on the core node, from a thread about driving a prompt box from an LLM: "The generatetext node is pretty naive and slow compared to a proper framework." That's the gap.

What you'd actually do with it

Three jobs, roughly in order of demand: rewriting a rough idea into a structured prompt; captioning an image with a vision-language CLIP; or generating text inside the graph. The reason to do it here rather than through Ollama is that the model is already downloaded and loaded. No second 8B file, no second process.

The tooltips point at encoders you already have: Qwen3 4B via Load CLIP with type lumina2, or Gemma 3 with type ltxv. Run Z-Image and you're one Load CLIP node from a local text writer.

How the fast path works

It reimplements nothing. It flips ComfyUI's own switches for the length of one generation and puts them back.

The node digs a transformer out of the CLIP, checks it generates through comfy.text_encoders.llama.BaseGenerate.generate, and confirms the layer stack has what the fixed-cache path needs. Then it probes the flash-decode kernel - available on this device? does it accept this model's head size? If both are yes, it sets fixed_kv on the model, adds the graph-capture flags where CUDA graphs are enabled and the patcher is dynamic, and restores the old values afterwards.

Where it can't, it says so and falls back to what core Generate Text does. The case you'll actually meet: prompt plus max_length running past the model's smallest sliding attention window, which drops the run back to core decode rather than generating garbage. The panel reports tokens per second and which decode ran.

The inputs that matter

clip and prompt are the two you always touch. The clip is a Load CLIP output; the prompt is a multiline string with dynamic-prompt support, so {a|b|c} works. max_length caps output at 512 tokens, which the tooltip translates as "about 380 words" - leave it and move on.

sampling_mode is the interesting one. Set off, the model takes the likeliest token every step and writes the same text every time. Set on and a settings group appears: temperature, top_k, top_p, min_p, repetition_penalty, presence_penalty and a seed - your reproducibility lever. Defaults are sane (temperature 0.7, repetition_penalty 1.05).

Then the optional row. image, video and audio feed a model that can see or hear - video takes an IMAGE batch, read at 24 fps and sampled at 1 fps. thinking lets a reasoning model deliberate first. use_default_template should stay on: it wraps your prompt in the model's own chat template, which is why these encoders read prompts like instructions rather than bags of tags. mtp is multi-token prediction for checkpoints carrying those heads; fast_decode is the one you'll flip off to compare against core.

One output, generated_text, a STRING - into a CLIP Text Encode's prompt, a Save Text node, or a preview. Prompt enhancement usually looks like Fast Generate Text → CLIP Text Encode → sampler.

Installing it

ComfyUI Manager, search WAS Node Suite v3, restart. Or:

cd ComfyUI/custom_nodes
git clone https://github.com/WASasquatch/was-node-suite-comfyui.git

The pack installs nothing and downloads nothing. This node needs ComfyUI 0.14.0 or newer - the flags and the kernel are ComfyUI-side - and Python 3.10+. The model comes from core Load CLIP, so if Qwen3-4B is already on disk for Z-Image you have no setup left.

Where people get burned

It ran on core decode and you want to know why. The panel names the reason: fast decode switched off on the node, a CLIP with no language model, a transformer generating through its own decoder, no fixed-cache path on the model, an unavailable kernel, a head size the kernel refuses, or the sliding-window case above. --disable-cuda-graphs kills the graph-capture half too.

You wired in something that isn't a Qwen/Gemma-family encoder, or left clip empty. Both fail loudly: the second names qwen_3_4b.safetensors to tell you which kind of CLIP it wants.

The text repeats itself. Small models loop. Nudge repetition_penalty up from 1.05, or drop temperature.

Expectations, honestly. thinking on a Qwen3 buys deliberation tokens and, when you're rewriting a prompt, more chance of the model's scratch-work leaking into your output - which is why the reasoning-model class barely registers in this corner of the ecosystem (docs/knowledge/llm-in-comfyui.md § 3). And mind the VRAM: a 4B encoder is nothing, but LTX-2's Gemma 3 12B is the 22.7 GB file behind most of that model's launch-day OOM errors. It's loaded for LTX-2 anyway; don't wire it into a graph that never needed it.

CategoryWAS Suite/Text

Inputs (11)

NameTypeDefaultDescription
clipCLIPA CLIP holding a language model, such as Qwen3 4B from Load CLIP with type `lumina2`, or Gemma 3 with type `ltxv`.
promptSTRINGWhat to ask the model, such as `Describe a lighthouse on a stormy night.`
max_lengthINT5121–32768The most tokens to write. A token is about three quarters of a word, so `512` is about 380 words.
sampling_modeCOMBO`on` draws each token at random within the limits below. `off` always takes the likeliest token and writes the same text every time.
imageoptIMAGEA picture to ask about, for a model that reads images such as Qwen 2.5 VL or Gemma 3.
videooptIMAGEVideo frames as an image batch, read as 24 fps and sampled at 1 fps.
audiooptAUDIOSound to ask about, for a model that hears it.
thinkingoptBOOLEANfalse`true` lets a model that reasons, such as Qwen3, think before answering.
use_default_templateoptBOOLEANtrue`true` wraps the prompt in the model's own chat template and system prompt.
mtpoptCOMBOautoMulti-token prediction for a checkpoint carrying those heads. `auto` picks the draft depth, `2` to `5` fixes it, `off` turns it off.
fast_decodeoptBOOLEANtrue`true` decodes on the fixed cache graph path where the model allows it; `false` runs exactly as core Generate Text.

Outputs (1)

NameTypeDescription
generated_textSTRINGWhat the model wrote, special tokens removed.