Gprompts
Another dynamic prompt node, designed to be easy to use and support wildcards
Nodes (6)
Your PNG can carry a note now — and one loader actually reads it back
{random|sequential} prompt text, minus the ceremony
Immich won't take your mp3, so this node builds a one-frame music video
Put your clips somewhere you'll actually find them
One image per run, no python loop required
Full Python f-string formatting inside ComfyUI
ComfyUI GPrompts Nodes
Introduction
This package provides custom nodes for ComfyUI that enhance prompt generation, string formatting, and saving images, video, and audio.
Nodes Overview
- GPrompts - Create dynamic prompts with random or sequential selection. Also supports wildcard files.
- Dynamic Prompts with Enhancer + Prompt Enhancer Loaders (GGUF / API, text-to-image / edit) + Enhancer Pair - GPrompts plus an LLM prompt enhancer (e.g. the Qwen-Image-2.1 prompt rewriters), local GGUF or any Ollama/OpenAI-compatible/Anthropic API, for text-to-image, image editing and MiniMax H3 video prompts, calling the LLM every run or once per batch.
- String Formatter - Build custom output strings from multiple inputs and system variables.
- Save Image With Notes - Save images with embedded workflow notes and metadata.
- Load Images From Folder - Load images from a folder one at a time, in order, or randomly, and read back the prompt they were created with.
- Save Image To Immich Server - Save images with embedded workflow notes to an Immich server.
- Save Video To Immich Server - Save videos to an Immich server.
- Save Audio To Immich Server - Save audio to an Immich server (as an mp4 with a cover image).
GPrompts Node
Description
This is another dynamic prompts node for ComfyUI. I found most of the ones out there to be either too complicated or too limiting, so I wrote my own.
Basics
Create a GPrompts node and connect its output to a CLIP node text input.
Format of a dynamic prompt:
{ cat | dog | jackalope } random selection
{{ green | yellow | red }} sequential selection
If you generate 4 images with a stop light showing {{ green | yellow | red }}
you will get a green light image, yellow, red, and then another green.
If you use a stop light showing { green | yellow | red }, each image will have a 33% chance of any color.
The sequential cycle starts at the first combination and starts over whenever the prompt text changes (or ComfyUI is restarted).
The normal delimiters are { and }, but you can change them to < and >, which is useful for JSON prompts:
< cat | dog > random selection
<< green | red >> sequential selection
Registers
{0 cat | dog } will store the chosen value in register 0, which can be referenced later with {0}.
{ and {{ have their own separate registers. Registers 0..9 are available.
a {0 green | blue} {{0 dog|monkey}} standing on top of a {0} {{0}}
will generate something like a green dog standing on top of a green dog.
Note: {{0|1|2}} (no space after the digit) is a normal options block, not a register.
Quotes and special characters inside blocks
Avoid using " and ' inside { } and {{ }} blocks. Double quotes in particular will stop a random block from being recognized, and quotes can break JSON prompts.
Two ways around this:
-
Keep the quotes outside the block. This works fine:
wearing a T shirt that says "{ hello | goodbye }" -
Use a Unicode character that looks like a quote. Copy and paste one of these into your options:
Instead of
"(double quote)| Character | Name | Code point | |-----------|------|------------| |
“| Left double quotation mark | U+201C | |”| Right double quotation mark | U+201D | |″| Double prime | U+2033 | |"| Fullwidth quotation mark | U+FF02 | |„| Double low-9 quotation mark | U+201E | |〃| Ditto mark | U+3003 |Instead of
'(single quote / apostrophe)| Character | Name | Code point | |-----------|------|------------| |
‘| Left single quotation mark | U+2018 | |’| Right single quotation mark | U+2019 | |′| Prime | U+2032 | |'| Fullwidth apostrophe | U+FF07 | |ʼ| Modifier letter apostrophe | U+02BC | |‚| Single low-9 quotation mark | U+201A |Example:
a sign that reads { “OPEN” | “CLOSED” | “BACK IN 5” } a { cat’s | dog’s } toy
If you need a literal | inside an option, escape it with a backslash: \|.
Wildcards
Wildcard files are either .txt or .json and go in comfyui/models/wildcards.
You can use a wildcard file with a list of options:
a woman with {{__hair_color__}} {{__hair_style__}} hair
This will use the contents of comfyui/models/wildcards/hair_color.txt and hair_style.txt.
Assuming the files are
blonde
red
brown
and
long
short
pixie
mohawk
you will have 12 combinations. Set ComfyUI to generate 12 images and you will see all combinations.
A wildcard reference to __hair__hairstyles__ will use the file models/wildcards/hair/hairstyles.txt (or .json).
JSON wildcard files
Instead of text you can use a JSON file. For example seasons.json:
Simple:
{ "doesnotmatter": ["summer", "winter", "fall", "spring"] }
Weighted:
{ "whatever": [ { "summer": 6 }, { "spring": 4 }, { "fall": 3 }, { "winter": 1 } ] }
Weights are only relevant to {} random selection. With random, the odds of getting a choice are weight / total weight, so for summer the odds are 6 out of 14.
For {{}} sequential you will get all 4 seasons.
TODO
- Add support for wildcard files that include other wildcard files.
Dynamic Prompts with Enhancer
Description
Same dynamic prompt syntax as GPrompts, plus an LLM that rewrites the prompt into the long, detailed form image models like Qwen-Image 2.1 work best with.
Which models it suits:
- Qwen-Image 2.1 - recommended; this is what it was built and tuned for, with the Qwen-Image prompt rewriters (see Getting the models).
- MiniMax H3 video - recommended; see Video prompts for MiniMax H3 below.
- Ideogram 4 - recommended, and close to required: Ideogram 4 is trained only on JSON captions and plain text works poorly. See JSON captions for Ideogram 4 below.
- FLUX.2 (especially Klein) - not recommended. FLUX, and Klein in particular, looks most
realistic with short, plain prompts; a detailed rewrite makes skin and surfaces glossy and
artificial. For FLUX, connect the
dynamic_promptoutput (your expansion, not rewritten) to the text encoder instead oftext, or use the plain Dynamic Prompts node.
The enhance setting picks how often the LLM runs:
- every run (default) - each run's expansion is sent to the LLM. One LLM call per image, and
every prompt is rewritten around its own values. With a seed above 0 the LLM gets
seed + run number, so a combination that comes round again gets a fresh (but reproducible) rewrite. - once, then substitute - only the first expansion goes to the LLM. For every later run the
node swaps that run's chosen values into the LLM's rewrite, so a whole batch costs one LLM call.
Example:
a {{tiki bar|beach hut}} with {red|amber} lanternsis rewritten once as "Cinematic photo of a tiki bar with amber lanterns at dusk", and the next run becomes "Cinematic photo of a beach hut with red lanterns at dusk". This only works while the LLM keeps the values as words that can be swapped. If it rewords or elaborates them ("add a dinosaur" becomes "a towering sauropod..."), later runs can't swap in their value; use every run for prompts like that.
With every run, prefetch_next hides most of the LLM's time: as soon as a run's prompt is out, the node expands the next run and starts rewriting it on a background thread, so the LLM works while the image renders, and the next run usually finds its prompt ready.
auto(default) - on for an API loader pointing at another machine or a hosted service, and for a GGUF on the CPU (gpu_layers= 0). Off for an API server on this machine (localhost, a blankbase_urlfor the local styles, or this machine's own name or address), which shares the GPU.on- also prefetch from a server on this machine (it may slow the image a little; a failure there only fails that rewrite).off- never.
A GGUF on the GPU never prefetches, not even with on. It runs inside the ComfyUI process, and
llama.cpp aborts the whole process on a GPU error, such as running out of memory while the image
model samples at the same time. That takes ComfyUI down (Fatal Python error: Aborted, core
dumped). To prefetch with a local model, run it on the CPU (gpu_layers 0; the vision projector
then runs on the CPU too) or behind a server such as Ollama or llama-server.
The next run's expansion is predictable for sequential {{ }} blocks and for random { } blocks
with a fixed seed. With the seed set to randomize and random blocks in the text, the prediction is
usually wrong: the background rewrite is dropped and that run calls the LLM itself (no gain, no
harm). After the last run of a batch one background rewrite goes unused, which on a paid API is one
wasted call.
In substitute mode, preserve_dynamic_words (on by default) asks the LLM to keep the chosen words verbatim. If it rewords one anyway ("cat" becomes "kitten"), that variation can't be swapped in; the console says so and that run keeps the first run's value.
LoRA trigger words go in trigger_words (just below the text), not in the text: an LLM
takes an unknown word like ohwx for a typo and drops or "corrects" it. The node adds them to the
start of the final prompt exactly as typed, after the LLM has run, in every mode:
ohwx woman, tkb_style + the LLM's "A woman in a red silk dress..." gives
ohwx woman, tkb_style, A woman in a red silk dress.... Separate several with commas. One the
prompt already contains is not added twice. Changing them never costs a new LLM call. Describe the
subject in plain words in the text ("a woman in a bar") so the LLM writes a normal description. The
start of the prompt is where most LoRAs expect their trigger, and where the image model pays most
attention.
Changing the text, delimiter style, enhance mode, any enhancer setting or a reference image starts
the batch over. The seed does not: as in Dynamic Prompts it only picks the random blocks, so a seed
set to randomize still steps through the combinations.
Outputs
text/computed_prompt- the final prompt for this rundynamic_prompt- this run's expansion before enhancement (with the trigger words in front too)seedtemplate- the prompt exactly as typed, dynamic blocks and allwh_ratio- aspect ratio recommended by the LLM (e.g.16:9), empty if nonewidth/height- an image size atwh_ratio(1:1 if none) with abouttarget_megapixelspixels (1.0 = 1024x1024-sized, 4.0 = Qwen-Image 2.1's native 2K), sides rounded to multiples of 16. Wire them into an Empty Latent Image for text-to-image.resolution- the side of a square with abouttarget_megapixelspixels (multiple of 32; 1.0 = 1024), for theresolutioninput of Text Encode Qwen Image 2.1 in edit workflows. For text-to-image, don't use that node's latent output: it is always square. Use width/height into an Empty Latent Image instead; the encode node'sresolutionthen has no effect.ratio_follow- for edits, the reference image whose shape the output keeps (e.g.image1), as chosen by the LLM; width/height follow that image's shape.
Getting the models
The Qwen-Image-2.1 prompt rewriters come as two separate models, one for text-to-image (PE-T2I)
and one for image editing (PE-I2I), each with its own system prompt. The edit model also needs
its vision projector (mmproj), the file that lets the LLM see images. Any quantization works
(Q4_K_M ~6 GB, Q6_K ~7.5 GB). For example, with the hf command from huggingface_hub, run from
the ComfyUI folder:
hf download pottokao/Qwen-Image-2.1-PE-T2I-Heretic-GGUF \
pe_t2i_heretic-Q4_K_M.gguf system_prompt.txt --local-dir models/LLM/pe_t2i
hf download pottokao/Qwen-Image-2.1-PE-I2I-Heretic-GGUF \
pe_i2i_heretic-Q4_K_M.gguf pe_i2i_heretic.mmproj-bf16.gguf system_prompt.txt \
--local-dir models/LLM/pe_i2i
or download the same files in a browser from the repos' Files tab. Keep each model in its own
subfolder: both ship a file called system_prompt.txt, and the edit one (~18 KB) is different from
the text-to-image one (~10 KB).
ComfyUI/models/LLM/
pe_t2i/ pe_t2i_heretic-Q4_K_M.gguf, system_prompt.txt
pe_i2i/ pe_i2i_heretic-Q4_K_M.gguf, pe_i2i_heretic.mmproj-bf16.gguf, system_prompt.txt
Other GGUF builds work too, but the mmproj must come from the same repo as its model. After adding files, refresh the browser page (or restart ComfyUI) so the node lists pick them up.
Prompt Enhancer Loader (GGUF)
Put the .gguf model and its system prompt file (.txt or .md) in ComfyUI/models/LLM/
(subfolders are fine), then pick them on the node. The model loads the first time a prompt is
enhanced; turn keep_loaded off to free its VRAM after each use, or set gpu_layers to 0 to run it
on the CPU and leave the GPU to the image model.
Note that gpu_layers defaults to -1: the whole LLM goes on the GPU and, with keep_loaded on,
stays there next to the image model. A 9B rewriter at Q4-Q6 holds roughly 6-8 GB of VRAM plus its
context. If the image model runs short of memory, turn keep_loaded off or set gpu_layers to 0
(slower rewrites, but no VRAM used).
For the Qwen-Image-2.1 prompt rewriters, use the system_prompt.txt that ships with the model; they
produce nothing useful without it. Keep enable_thinking on, and use presence_penalty ~1.5 for the
text-to-image rewriter.
Requirements
The loader runs the model in-process with the llama_cpp package
(llama-cpp-python), the same way
ComfyUI-Prompt-Enhancer does. It is
optional and not installed automatically (see requirements-local-gguf.txt), because it has to match
your CUDA version. Install it into ComfyUI's own Python.
Easiest, no compiling: download the prebuilt CUDA wheel for your CUDA and Python version from
JamePeng/llama-cpp-python releases
(the file name carries both, e.g. +cu128 and cp312) and python -m pip install the .whl.
Or build it yourself:
CMAKE_ARGS="-DGGML_CUDA=on" python -m pip install -U llama-cpp-python --no-cache-dir
Qwen3.5-based models, including the Qwen-Image-2.1 prompt rewriters, need llama-cpp-python 0.3.35+ (PyPI) or the JamePeng fork 0.3.47+.
Thinking and plan_tokens
With enable_thinking on, the model plans before it answers. That plan is where the time goes, so
plan_tokens caps it: when the plan reaches that length it is handed back to the model, closed, and
the model writes the answer from it. -1 removes the cap (slowest, most detail). With thinking off the
answer is written directly (fastest, shallower prompts).
Prompt Enhancer Loader (API)
The same enhancer, but the LLM runs behind an HTTP API: a local or LAN server (Ollama, llama-server, vLLM, SGLang, LM Studio) or a hosted one (OpenAI, OpenRouter, DashScope, Anthropic, or anything OpenAI-compatible). Running the LLM on another machine leaves all your GPU memory to the image model. No extra Python packages are needed.
Pick the request style with api:
| api | endpoint | use it for |
|---|---|---|
| OpenAI-compatible chat | /chat/completions | almost everything |
| OpenAI-compatible completions (raw prompt) | /completions | Qwen-family models on vLLM, SGLang, llama-server, LM Studio |
| Ollama chat | /api/chat | any Ollama model |
| Ollama generate (raw prompt) | /api/generate | Qwen-family models on Ollama |
| Anthropic messages | /v1/messages | Claude |
The two raw prompt styles send exactly the prompt the GGUF loader builds (think prefill,
plan_tokens, early stop), so a Qwen-Image prompt rewriter behaves the same as it does locally. Chat
styles let the server apply the model's own template; thinking is switched on or off with whatever
field the server understands (thinking_field, auto-detected by default).
Servers differ in which parameters they accept. If a request is rejected because of a parameter
(top_k, min_p, seed, max_tokens, a thinking field...), the node drops or renames it, retries,
and remembers that for the server and model. extra_json adds anything provider-specific to the
request body.
API keys go in Settings > Gadzoinks > LLM, never on the node (node values are saved into workflows and image metadata):
- LLM API key (default) - used when
api_key_nameis blank. Leave it empty for local servers. - LLM API keys, named -
openrouter=sk-or-...; dashscope=sk-...; select one withapi_key_name(e.g.openrouter). api_key_namealso acceptsenv:VARIABLE(read an environment variable) andnone.
For Ollama, set context_length (Ollama's small default context silently truncates the 10 KB
Qwen-Image system prompt) and use keep_alive = 0 to unload the model right after each call.
Quick start: a Qwen-Image rewriter on Ollama
Useful for running the LLM on another machine (a Mac, a second PC) so the ComfyUI GPU is left to the image model.
- On the LLM machine, import the GGUF into Ollama. No template or system prompt is needed in the
Modelfile: the node sends both.
echo 'FROM ./pe_t2i_heretic-Q4_K_M.gguf' > Modelfile ollama create qwen-pe-t2i -f Modelfile ollama list - If ComfyUI is on a different machine, make Ollama listen on the network:
OLLAMA_HOST=0.0.0.0 ollama serve, or for the macOS applaunchctl setenv OLLAMA_HOST 0.0.0.0and restart the app. Check from the ComfyUI machine:curl http://<llm-host>:11434/api/tagsshould list the model. Ollama has no authentication; only do this on a network you trust. - Keep the matching
system_prompt.txtinmodels/LLMon the ComfyUI machine. - On Prompt Enhancer Loader (API) set
api=Ollama generate (raw prompt),base_url= blank (same machine) orhttp://<llm-host>:11434,model=qwen-pe-t2i,system_prompt= the T2Isystem_prompt.txt,context_length= 16384. Leaveapi_key_nameblank. - Queue. The console shows
GPromptsEnhanced: Ollama generate (raw prompt) -> ... model=qwen-pe-t2iand then... finish=earlywhen the answer is complete.
If ollama create rejects the GGUF's architecture, update Ollama. Or serve the GGUF with
llama.cpp's own server instead (llama-server -m pe_t2i_heretic-Q4_K_M.gguf -c 16384 --host 0.0.0.0 --port 8080) and use api = OpenAI-compatible completions (raw prompt) with base_url =
http://<llm-host>:8080/v1.
For edits over an API, the server must see images: use a chat style on the edit loader and a
server that serves the PE-I2I model together with its mmproj (e.g. llama-server -m pe_i2i_heretic-Q4_K_M.gguf --mmproj pe_i2i_heretic.mmproj-bf16.gguf ...).
Image editing
Connect reference images to Dynamic Prompts with Enhancer: its image_1 socket grows a new one
each time you connect an image (up to 16, like Text Encode Qwen Image 2.1). Refer to them in your text
as <image1>, <image2>, ... When images are connected, the node uses an edit loader; without
them, a text-to-image loader.
| Loader | Use |
|---|---|
| Prompt Enhancer Loader (GGUF) / (API) | text-to-image |
| Prompt Enhancer Loader (GGUF, edit) / (API, edit) | image edit (the LLM sees the images) |
| Enhancer Pair | one text-to-image + one edit loader behind a single enhancer output |
An edit model needs its own files: for Qwen-Image-2.1 the PE-I2I GGUF, its mmproj vision
projector (also in models/LLM), and its own ~18 KB system prompt - not the text-to-image ones. Use
presence_penalty 0 for it. Images are downscaled for the LLM (llm_image_megapixels, ~512 tokens per
image at 0.5); the encode node still receives the originals. The API edit loader offers only the chat
styles, since raw prompts can't carry images; the server must serve a vision model (llama-server with
--mmproj, an Ollama vision model, or a hosted one).
Wiring an edit: each Load Image goes to two places, in the same order:
image_1,image_2, ... on Dynamic Prompts with Enhancer (so the rewriter sees them), andimage_1,image_2, ... on Text Encode Qwen Image 2.1, with the VAE connected (so the image model edits them).
Then text -> the encode node's prompt, resolution -> its resolution, and use the encode node's
own latent output for the sampler (it is sized from the first reference; any other size shifts the
edit). In once, then substitute mode the images stay fixed for the batch while the dynamic
blocks vary the instruction, so the batch costs one LLM call; changing an image starts a new one.
Video prompts for MiniMax H3
MiniMax H3 (Hailuo 3) generates video with sound, and expects a structured prompt: labelled
sections for the picture timeline, the soundscape and the music, shot timestamps, camera terms,
<Picture 1>-style reference labels and <d>[English] ...</d> dialogue tags. The folder
system_prompts/minimax_h3/ in this pack has three system prompts, written from MiniMax's official
prompt-writing guide, that turn a short request into that format:
| System prompt | For | Loader | H3 node |
|---|---|---|---|
| h3_t2va_system_prompt.txt | text-to-video | Prompt Enhancer Loader (API) or (GGUF) | H3 Text to Video |
| h3_frames_system_prompt.txt | first frame, last frame, or both | ... (API, edit) or (GGUF, edit) | H3 First-Last-Frame to Video |
| h3_reference_system_prompt.txt | reference-to-video (up to 9 images, plus videos/audio) | ... (API, edit) or (GGUF, edit) | H3 Reference to Video |
Setup
- Copy the three files to
ComfyUI/models/LLM/h3/and refresh the browser page. - Pick one on the loader's
system_prompt. These prompts ask for plain text, not JSON: the node uses the whole answer, andwh_ratio/width/heightdon't apply (the H3 node sets the size). - Connect
textto the H3 node's prompt input.
Wiring
- Frames and reference images go to both nodes, in the same order:
image_1,image_2, ... on Dynamic Prompts with Enhancer (so the LLM sees them), and the first/last frame or reference image inputs on the H3 node. For frames: one image = first frame (say "last frame" in the text if it is the last one), two images = first and last. - Reference videos and audio go to the H3 node only; mention them in your text ("the dance from video 1", "her voice from audio 1"). The LLM can't watch or hear them and works from your words.
- Duration: the prompt's shot timings must fit the video. Write the length in your text if it
isn't the default ("10 seconds: ..."); the LLM rounds it up to the next length H3 renders (table
below). Set the matching
length(frames) on the local H3 latent node; the online H3 nodes take whole seconds, so set the same number there (a fraction of a second difference doesn't matter). - Leave
trigger_wordsempty: H3's prompt must start with its own first line or section. - Use
enhance= every run; thinking can stay off. - Temperature: with an H3 system prompt and the loader's
temperatureleft at its default 1.0, the node samples at 0.3, which keeps the model on the format. Any other value you set is used as it is (Claude with thinking on always runs at 1.0). - Format check: after each rewrite the node checks the answer against MiniMax's format (sections and their order, the first line, shot numbers and cut times, shots of at least 1.2 s, cut phrases, dialogue tags, labels, music in the soundscape) and logs a warning listing anything off. It never changes the prompt; read the warning and re-queue or edit the text.
- 2nd try on error (loader toggle, default off): when the check finds problems, the node sends
them back to the LLM with its answer and asks for a corrected prompt, once. It keeps whichever
answer has fewer problems (the first on a tie) and logs
2nd try: 3 -> 0 problem(s); using the second answer. It costs a second LLM call, only when the first answer fails; on Claude the system prompt and images come from the prompt cache. It applies to the H3 and Ideogram 4 system prompts.
Local H3 renders 17k+5 frames at 24 fps (ComfyUI snaps the latent length up to that), and was
trained on 124-362 frames. The lengths it can make:
| Seconds | 5.17 | 5.88 | 6.58 | 7.29 | 8.00 | 8.71 | 9.42 | 10.13 | 10.83 | 11.54 | 12.25 | 12.96 | 13.67 | 14.38 | 15.08 | |---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---| | Frames | 124 | 141 | 158 | 175 | 192 | 209 | 226 | 243 | 260 | 277 | 294 | 311 | 328 | 345 | 362 |
With no length in your text the prompt is written for 5.17 s (124 frames, the H3 node's default).
Which LLM
- Claude (Anthropic messages, e.g.
claude-haiku-4-5-20251001) follows the format reliably and sees your images; roughly half a cent to a cent per prompt. - Local: a general vision model through the GGUF edit loader, e.g. a Qwen3.5-9B instruct GGUF with
its mmproj. Set
keep_loadedoff so it leaves the GPU before H3 loads, or run it on another machine (llama-server with--mmproj) behind the API edit loader. Small models handle text-to-video and frames well but may slip on the long six-section reference format. - MiniMax's own enhancer is also built into ComfyUI: MiniMax H3 Context IR (Prompt Enhancer). It writes the same format and can also watch reference videos and hear audio, but runs on Comfy credits (about $0.05-0.11 per call).
H3 itself runs either as ComfyUI's built-in MiniMax H3 partner nodes (online, Comfy account and credits) or locally with the open weights and ComfyUI's native H3 workflows (see the ComfyUI docs; it is a 33B model, so plan on a 24 GB GPU with the INT8/Q5 files, or 16 GB with Q3/Q4/INT4).
JSON captions for Ideogram 4
Ideogram 4 was trained only on structured JSON captions: a short high_level_description, then a
compositional_deconstruction with the background and a list of elements, each an object or a
piece of text with an optional bounding box (bbox, [y1, x1, y2, x2] on a 0-1000 grid). Plain
text prompts give poor results and trip its safety filter more often. The folder
system_prompts/ideogram4/ has ideogram4_system_prompt.txt, which turns a short idea into that
JSON. It is Ideogram's own open-source magic prompt (Apache 2.0, see NOTICE in that folder),
adapted to this node.
Setup
- Copy
ideogram4_system_prompt.txttoComfyUI/models/LLM/ideogram4/and refresh the page. - Pick it on a text-to-image loader (API or GGUF). Thinking off;
max_new_tokens4096 or more (a busy poster can be 2,000+ tokens of JSON);context_length16384 for local models. - Connect
textto the Ideogram 4 prompt input, andwidth/heightto the latent size. - Set
delimiter_styletoangle < >if you write JSON intext(see below).
Your input: a plain idea or a JSON draft
- A plain idea ("a ramen stall in a Tokyo alley at night, neon sign reading 'ラーメン'") works with either delimiter style.
- You can also write all or part of the caption yourself as JSON, with dynamic choices inside it.
Use
angle < >delimiters so the choices can't be confused with the JSON's own braces, and a register when one choice appears in several fields, so they agree:
The LLM treats the draft as the brief: it keeps your elements, text, boxes and colors, fills in what is missing and fixes the schema.{"high_level_description":"A <0 tabby cat|black dog> on a <red|green> sofa.", "compositional_deconstruction":{"background":"Sunlit living room.", "elements":[{"type":"obj","bbox":[350,250,900,700],"desc":"The <0>, curled up asleep."}]}} - If your JSON is already a complete caption and you don't want it rewritten, skip the LLM: use the
plain Dynamic Prompts node with
angle < >and connect it straight to the prompt input.
What the node does with the answer
- Checks it is valid JSON and parses it (code fences, a preamble or
<think>are dropped), puts the keys in Ideogram's required order, and passes it on as one line with non-ASCII text kept. - The LLM also picks an aspect ratio (from words like
16:9,poster,widescreenin your text, or by itself). The node removes it from the caption and sends it towh_ratio, sowidth/heightmatch the layout the bounding boxes were planned for. Use those outputs, or say the ratio in your text if your latent size is fixed. trigger_wordsgo to the front ofhigh_level_description, inside the JSON.- Checks the caption against Ideogram's schema (bounding boxes, colors, element types) and logs a warning if anything is off; with 2nd try on error on, it asks the LLM to fix it once.
Which LLM
- Claude (e.g.
claude-haiku-4-5-20251001, or Sonnet/Opus): Ideogram tested this prompt with Claude Opus. It is a long system prompt (about 7,000 tokens), so on Claude it comes from the prompt cache after the first call. - Local: a general instruct model of 9B or larger (e.g. Qwen3.5-9B) can do it, but follows the many rules less closely; turn on 2nd try on error. Not the Qwen-Image PE rewriters.
Notes
- Ideogram's own pipeline removes the bounding boxes before rendering by default, keeping only the descriptions. This node keeps them, for layout control. If subjects come out duplicated or squeezed into their boxes, check that the latent has the caption's aspect ratio.
- The Ideogram 4 weights are under a non-commercial licence (the magic prompt itself is Apache 2.0).
Warnings
- every run costs one LLM call per image. With thinking on, that can take longer than the image
itself; use
prefetch_next(API or CPU), a lowerplan_tokens, or thinking off. On a paid API, every image is a paid request. - once, then substitute is only safe for simple swaps (colours, plain nouns). If the LLM
rewrites a value into something else, later runs keep the first run's word; the console prints
'<word>' not found in the enhanced prompt, so '<new word>' was not appliedwhen that happens. - Local GGUF uses the GPU by default (
gpu_layers-1) and competes with the image model for VRAM. It is never run in the background (seeprefetch_next): a GPU error there would abort ComfyUI. - prefetch_next makes one extra call per batch (the rewrite for a run that never comes). Set it
to
offfor paid APIs if that matters. - Your prompts leave the machine when the API loader points at a hosted service (OpenAI, OpenRouter, Anthropic, DashScope, ...), and reference images do too with the edit loader.
- Node values are saved into workflows and image metadata. That is why API keys only go in
Settings > Gadzoinks > LLM; the node refuses
extra_jsonorapi_key_namevalues that look like a key. - The H3 video prompts and the image prompts are not interchangeable either: an H3 system prompt with an image model (or the other way round) gives a useless prompt.
- The edit and text-to-image files are not interchangeable. The PE-I2I model with the T2I system prompt (or the other way round) gives poor or broken prompts.
Troubleshooting
| Message or symptom | Cause and fix |
|---|---|
| mmproj must be the vision projector file, not the model itself (shown under model, mmproj and system_prompt) | There is no mmproj file in models/LLM, so the list fell back to all GGUFs. Download the model's ...mmproj...gguf (see Getting the models), refresh, and select it. ComfyUI repeats the one error under every input it checks. |
| GGUF model / mmproj / System prompt file not found in models/LLM | The file was moved or renamed; refresh the page and pick it again. |
| The GGUF prompt enhancer needs the llama_cpp package | Install llama-cpp-python (see Requirements), or use the API loader. |
| Reference images are connected, but the enhancer is a text-to-image loader (or the reverse) | Use the matching loader, or an Enhancer Pair with both. |
| API: connection refused or timeout | The server isn't running or isn't listening on the network (OLLAMA_HOST), or base_url is wrong. |
| API: HTTP 404 | model doesn't match the server's name for it (ollama list). |
| API: HTTP 401/403 | Missing or wrong key in Settings > Gadzoinks > LLM, or the wrong api_key_name. |
| Output ignores the rewriting rules, or no JSON / no wh_ratio | No system prompt selected, the wrong one, or (Ollama) context_length too small for it. |
| '<word>' not found in the enhanced prompt, so ... was not applied | once, then substitute mode and the LLM reworded that value. Switch enhance to every run. |
| ComfyUI dies with Fatal Python error: Aborted in llama_cpp ... ggml_abort | llama.cpp hit a GPU error, usually out of VRAM next to the image model. Turn keep_loaded off, use a smaller quant or context_length, or run the LLM on the CPU (gpu_layers 0) or on another machine. |
| Very slow rewrites | Thinking is on with no cap: lower plan_tokens or turn thinking off. Check that a GGUF is really on the GPU (gpu_layers -1) or that ollama ps shows GPU. |
String Formatter
Description
Builds an output string from supplied inputs and from system variables.
For example, if you connect prompt (or computed_prompt) to A and seed to B, then the format string generating $a with seed $b on $hostname will generate a string like generating a smiling cat with seed 12345 on hal2000.
📝 System Variables Reference
This node provides access to various system variables that can be used in your workflows. Below is a complete list of available variables:
📅 Date & Time Variables
| Variable | Description | Example |
|----------|-------------|---------|
| datetime | Full date and time | 2024-01-15 14:30:25 |
| date | Current date | 2024-01-15 |
| time | Current time | 14:30:25 |
| time_24h | 24-hour format time | 14:30 |
| time_12h | 12-hour format time | 02:30 PM |
| iso_datetime | ISO format datetime | 2024-01-15T14:30:25.123456 |
| timestamp | Unix timestamp (seconds) | 1705329025 |
| timestamp_ms | Unix timestamp (milliseconds) | 1705329025123 |
| year | Full year | 2024 |
| year_short | Short year | 24 |
| month | Full month name | January |
| month_num | Month number | 01 |
| day | Day of month | 15 |
| day_num | Day number (alias for day) | 15 |
| hour | Hour | 14 |
| minute | Minute | 30 |
| second | Second | 25 |
| am_pm | AM/PM indicator | PM |
| weekday | Full weekday name | Monday |
| weekday_short | Short weekday name | Mon |
🎲 Random & Unique Identifiers
| Variable | Description | Example |
|----------|-------------|---------|
| uuid | Full UUID v4 | 123e4567-e89b-12d3-a456-426614174000 |
| uuid_short | First 8 chars of UUID | 123e4567 |
| random_hex | Random hex string (32-bit) | a1b2c3d4 |
| random_int | Random 4-digit number | 7352 |
| counter | Sequential counter (6-digit, increments per execution) | 000042 |
| batch_id | Batch identifier (last 8 digits of timestamp_ms) | 90251234 |
🗂️ Path & Directory Variables
| Variable | Description | Example |
|----------|-------------|---------|
| date_path | Date formatted as path | 2024/01/15 |
| datetime_path | Datetime formatted as path | 20240115_143025 |
| cwd | Current working directory | /path/to/comfyui |
| model_dir | ComfyUI models directory | /path/to/models |
| input_dir | ComfyUI input directory | /path/to/input |
| output_dir | ComfyUI output directory | /path/to/output |
| temp_dir | ComfyUI temp directory | /path/to/temp |
💻 System Information
| Variable | Description | Example |
|----------|-------------|---------|
| hostname | Computer hostname | my-workstation |
| node | Network node name | my-workstation |
| os | Operating system with release | Windows 10 |
| system | System name | Windows |
| release | System release | 10.0.19045 |
| platform | Full platform info | Windows-10-10.0.19045 |
| machine | Machine type | AMD64 |
| processor | Processor info | Intel64 Family 6 Model 158 |
| architecture | System architecture | 64bit |
| cpu_count | Number of CPU cores | 16 |
| pid | Process ID | 12345 |
| python_version | Python version | 3.10.12 |
| user | Current username | username |
🎮 GPU Information (if available)
| Variable | Description | Example |
|----------|-------------|---------|
| cuda_available | CUDA availability | True / False |
| gpu_name | GPU device name | NVIDIA GeForce RTX 4090 |
| gpu_count | Number of GPUs detected | 1 |
Save Image With Notes
Description
This node modifies a copy of your workflow, adding a Note node to the new workflow that is then saved inside the image.
You can add your own text with the notes input,
or wire computed_prompt from GPrompts, which creates a Note and saves the computed prompt in the image metadata.
Note: This node uses the standard ComfyUI Save Image node to do the actual saving.
Load Images From Folder
Description
Loads images from a folder on the ComfyUI server. It can load a single image, step through the folder one image per run, or pick images at random.
It can also read the image metadata. If an image was saved with Save Image With Notes or Save Image To Immich Server with a computed prompt, the prompt output returns that prompt, so you can re-run or vary old generations.
Supported file types: png, jpg, jpeg, bmp, tiff, webp. Files are sorted by path, and EXIF orientation is applied automatically.
Credit: This node is forked from the Load Image Batch node in WAS Node Suite by WASasquatch. Thanks for the original work.
Node Settings
- mode:
single_image: loads the image atindex. Ifindexis larger than the number of images, it wraps around.incremental_image: loads the next image each time the workflow runs, and goes back to the first image after the last one. The position resets when ComfyUI restarts.random: picks a random image based onseed.
- seed: seed used by
randommode. - index: image number to load in
single_imagemode (starts at 0). - path: folder to load images from.
- pattern: filename pattern, default
*. For example*.pngfor PNG files only, or**/*to include subfolders. - allow_RGBA_output: if
false, images with transparency are converted to RGB. - filename_text_extension: if
true, thefilename_textoutput includes the file extension. - load_exif: if
true, reads metadata from the image so thepromptoutput can be filled in.
Outputs
- image: the loaded image.
- filename_text: the file name.
- width / height: image size in pixels.
- prompt: the computed prompt stored in the image, or empty if there isn't one.
Immich Nodes
Save images, video, and audio to an Immich server, and load images back from it: https://immich.app
Configuration (shared by all Immich nodes)
Create an API Key in your Immich server.
Install the node in ComfyUI and go to Settings. In Settings look for the Gadzoinks section. Enter the API Key, the Hostname, and the Port.
- Save to Disk: if disabled, the file is deleted from the ComfyUI server file system after uploading to Immich.
- Default Album: album to use if none is specified in the node.
- Default Tags: these tags are combined with the tags in the node.
Settings are remembered on the ComfyUI server, so saving keeps working after a restart or if the browser is disconnected.
To load with Load Image From Immich, the API key also needs album.read, tag.read,
asset.read, asset.download and asset.view (or simply all).
Save Image To Immich Server
Description
This node modifies a copy of your workflow, adding a Note node to the new workflow that is then saved inside the image.
You can add your own text with the notes input,
or wire computed_prompt from GPrompts, which creates a Note and saves the computed prompt in the image metadata.
Supports adding images to albums, and adding tags.
Node Settings
- notes: takes a string and creates a Note node that is added to the workflow saved with the image. Often used with the String Formatter node.
- computed_prompt: ignore, will probably be removed.
- album: add the image to this album. The album is created if it does not exist.
- tags: comma separated tags, merged with the Default Tags from Settings.
- save_also: if enabled, the image is saved as normal with ComfyUI. If disabled, the image on ComfyUI is deleted after upload.
Save Video To Immich Server
Description
Saves a video generated in ComfyUI and uploads it to your Immich server. Uses the same Gadzoinks settings as the image node, and supports albums, tags, and optionally keeping a copy on the ComfyUI server.
Save Audio To Immich Server
Description
Saves audio generated in ComfyUI and uploads it to your Immich server.
Immich only handles photos and videos, so it does not support audio files such as mp3. To get around this, the audio is saved as an mp4 video with a still cover image, which Immich can store and play.
- Cover image: optionally connect your own image to use as the cover. If no image is connected, a default cover is used.
- Supports albums, tags, and optionally keeping a copy on the ComfyUI server, the same as the image node.
Requirement: ffmpeg
ffmpeg must be installed and available on the system PATH of the machine running ComfyUI, or this node will not work.
- Windows:
winget install ffmpeg - macOS:
brew install ffmpeg - Ubuntu/Debian:
sudo apt install ffmpeg
You can check that it is installed by running ffmpeg -version in a terminal.
Load Image From Immich
Steps through the Immich images that match your filters, one image per run, like Load Images Batch does for a folder. Everything happens on the ComfyUI server, so it works from any browser and nothing is uploaded.
Filters: set any combination; an image must pass every filter that is set.
- album: only images in this album. Blank = any.
- tag: only images with this tag. Blank = any. For a nested tag use the full name, e.g.
Trips/Japan. - favorites_only: only images marked as favorite.
- min_rating: this many stars or more:
2+takes 2, 3, 4 and 5 stars. Unrated images count as 0, so any rating filter leaves them out.
Example: album Morocco, tag Canon, favorites_only on, min_rating 2+ gives the favorite
Canon shots in the Morocco album rated 2 stars or more.
Choosing names: the ▾ choose album and ▾ choose tag buttons at the bottom of the node
list the albums and tags in your Immich (type to filter the list). Picking one fills the box;
(any) clears it; ↻ refresh list reads them again from Immich (they are otherwise kept for a
minute). You can also type a name, or connect a text output to the box. Names are not
case-sensitive; one that isn't found gives an error listing the albums or tags there are.
- index: which image, oldest first by date taken; wraps around. Leave the control under it on
incrementand queue as many runs as thecountoutput says. Videos are skipped. - The list is read from Immich when the index is 0 and kept for the rest of the batch, so photos added to the album mid-batch don't shift the numbering. They are picked up at the next run from 0.
Outputs
image,mask(from transparency, as Load Image makes it),width,height.prompt,negative_prompt: read from the original file. In order: the prompt the Gadzoinks save nodes store (computed_prompt, trigger words and LLM rewrite included), the prompt traced through the ComfyUI workflow in the file (through the sampler's positive and negative inputs), or A1111/Forgeparameters(PNG text or JPEG EXIF). Empty for photos with no prompt in them.description: the description in Immich. For camera photos this is often the useful text; wire it into Dynamic Prompts with Enhancer to describe each photo.filename,asset_id,index,count(how many images match), andmetadata(JSON: the filters, date taken, people, tags, favorite, rating, and where the prompt came from).
File formats: the original is downloaded. HEIC photos need pillow-heif in ComfyUI's Python
(python -m pip install pillow-heif); without it, and for camera RAW files, the node uses Immich's
preview JPEG instead (about 1440 px, no prompt metadata) and says so in the log.
Immich versions: it uses the album and tag search filters that Immich 3.2 deprecated but still accepts, and falls back to the album's own asset list on older servers.
** Immich support ** All of these nodes work with standard Immich, but I have my own fork of Immich with extra features such as the ability to see the prompt and metadata of an Image, and the ability to search for text in a prompt ( find all images of dragons ) The installation is still rough, https://github.com/neal3000/immich_gadzoinks/tree/immich_gadzoinks
Logging and debugging
All nodes in this pack write to the ComfyUI console: the terminal window ComfyUI runs in, and the Logs tab in ComfyUI's bottom panel (terminal icon). Messages are tagged by level:
| Level | What you see | Shown by default |
|---|---|---|
| [INFO] | one short line per action: which LLM is called, prompt cache use, speed | yes |
| [WARNING] | something went wrong but the run continued: an LLM answer that wasn't the expected JSON, an HTTP retry, a failed Immich upload step | yes |
| [ERROR] | the run stopped; ComfyUI shows the traceback | yes |
| [DEBUG] | step-by-step detail: expansions, parsing, prefetch, settings sync, Immich upload steps | no - turn on below |
API keys are never written to the log, not even in debug output.
Turning on debug logging
Pick one of these.
1. In ComfyUI's settings (easiest)
- Open Settings (gear icon).
- Go to Gadzoinks > Debug.
- Switch on Debug logging (console).
It takes effect immediately, no restart needed. ComfyUI remembers the setting; after a restart it applies again as soon as the ComfyUI page is open in a browser. Switch it off the same way when you're done: debug output is verbose.
2. With an environment variable (from startup, and without a browser)
Set GPROMPTS_LOG=DEBUG before starting ComfyUI. It overrides the setting above, and is useful for
problems during startup or for headless/API use.
- Linux / macOS
GPROMPTS_LOG=DEBUG python main.py - Windows, Command Prompt
set GPROMPTS_LOG=DEBUG python main.py - Windows, PowerShell
$env:GPROMPTS_LOG = "DEBUG" python main.py - Windows portable build: edit
run_nvidia_gpu.bat(orrun_cpu.bat) and add the lineset GPROMPTS_LOG=DEBUGabove the line that starts ComfyUI.
Other values: INFO (the default), WARNING (only warnings and errors from this pack).
3. With ComfyUI's own --verbose flag
python main.py --verbose DEBUG turns on debug output for all of ComfyUI and every node pack, not
just this one. It works, but expect a lot of output.
Saving the output to a file
- Linux / macOS:
GPROMPTS_LOG=DEBUG python main.py 2>&1 | tee comfyui-debug.log - Windows:
python main.py > comfyui-debug.log 2>&1(the console then stays empty; open the file in a text editor) - Or copy it from the Logs tab in ComfyUI.
Reporting a problem
- Turn on debug logging (option 1 is enough).
- Reproduce the problem: queue the workflow again.
- Copy the console output from just before you queued until the error or wrong result.
- Include it in your issue, along with which nodes and loader settings you used.
For developers
Use the pack's logger rather than print:
from .common import get_logger
log = get_logger("mypart") # appears as gprompts.mypart
log.debug("detail %s", value) # %-style: the text is only built when debug is on
log.info("one line per action")
log.warning("recovered from a problem")
dprint(...) from common still works and logs at debug level.