Nodes/Krea2 Harness/Krea2 Prompt
ComfyUI Node

Krea2 Prompt

One Text Box, and a System Prompt You Never Get to See

By ANe5s·Created 30 days ago·Updated 2 days ago· 6
Krea2 Prompt
    • Merged Prompt
    • Main Prompt
    ◄main_prompt►

    Here's the thing nobody tells you when you download Krea 2. On Krea's website, you type six words and get a gorgeous image, because Krea runs its own prompt expansion on the server. The open weights don't come with that. Locally you get exactly what you typed, and the model - text-encoded by Qwen3-VL, built for long natural-language descriptions - does its best with a sentence fragment. That's not a bug in your workflow. It's the gap between the hosted product and the thing you downloaded, and Krea has admitted the open checkpoint also went through an alignment pass the API version didn't.

    This pack's answer is to run the expansion yourself, in the graph, with a small vision-language model. Krea2 Prompt is the front door of that pipeline. It's also the least interesting node in the pack, which is rather the point: it has one editable field and produces a wall of text you're not supposed to read.

    What it does

    You type your main prompt. The node wraps it in a fixed, versioned, role-separated ChatML system prompt ("Stage 1 V257" - the author says he iterated to version 262 and shipped 257) that is stored in the shipped Python as a Base64 constant, then tags an assistant turn onto the end. That's the "encoding/packaging, not cryptographic secrecy" note in the source, and it's accurate: the point is that you can't casually edit a thousand words of carefully-tuned instructions and break the pipeline.

    What's in that system prompt, roughly: identify the subject, mood and visual intent; pick one of two or three coherent readings rather than listing them; expand only what can actually be seen; preserve every explicit hard anchor (subject count, named identity, pose, gaze, camera angle, focus, requested text); don't invent a protagonist, animal, vehicle or location; return one natural-English paragraph of about 100–170 words with no headings, labels, JSON or markdown. There's also a patch appended at the end instructing the model to translate non-English input into English before expanding. So yes, a Chinese or Japanese prompt gets translated on the way through - that's deliberate, not a stray artifact.

    Inputs and outputs

    One input, two outputs, and both outputs matter.

    • main_prompt - your prompt, multiline. That's it. There are no hidden knobs.
    • Merged Prompt - the full role-separated transcript. This goes into the prompt input of a core TextGenerate node, along with Qwen3-VL loaded as the text encoder. It is a conversation, chat markers and all.
    • Main Prompt - your original text, unchanged. Wire it to original_main_prompt on Krea2 Prompt Harness. It's your fallback if the model misbehaves, and Stage 2 treats it as the authoritative "source ledger" for scene facts.

    The trap in the first thirty seconds

    Do not run Merged Prompt into a CLIPTextEncode. It is not conditioning text - it's a system-user-assistant transcript with <|im_start|> markers, and it will produce mush plus a confusing positive. The only thing it feeds is TextGenerate.

    The second trap is not wiring the Main Prompt output. That output looks decorative and isn't: everything downstream leans on it. Skip it and the harness nodes have nothing to fall back to when the 4B VLM answers as an assistant instead of a prompt engineer, which is the single most common way this pipeline fails.

    Installing it

    ComfyUI Manager, search the pack title Krea2 Harness. Manually:

    cd ComfyUI/custom_nodes
    git clone https://github.com/ANe5s/ComfyUI-Krea2-Harness.git
    

    Fully restart ComfyUI afterwards - a browser refresh won't reload Python nodes. The nodes appear under Nodes → ANe5s Nodes → Krea2. No extra Python dependencies, Python 3.10+, ComfyUI 0.3.0+.

    What you do need is models, and this pack ships none of them. The example text-to-image workflow expects a Krea 2 Turbo checkpoint in models/diffusion_models/, qwen3vl_4b_bf16.safetensors in models/text_encoders/, and qwen_image_vae.safetensors in models/vae/ - all of it from the Comfy-Org Krea 2 repo, plus the LoRAs the workflow references. The pack's example JSONs are in the repo's examples/ folder and each one carries a MarkdownNote with the download links and directories; drag the file onto the canvas and read it before you start fixing loaders.

    One expectation to set: your first queue now runs a language model before it runs the image model. On the 4B Qwen3-VL that's real seconds on every generation, and it's cached nowhere. That's the price of expansion, and it's why this node exists at all.

    CategoryANe5s Nodes/Krea2

    Inputs (1)

    NameTypeDefaultDescription
    main_promptSTRINGOriginal main prompt that the user can edit.

    Outputs (2)

    NameTypeDescription
    Merged PromptSTRINGRole-separated Stage 1 V257 prompt sent to the text-generation node.
    Main PromptSTRINGOriginal main prompt passed to source-ledger and fallback roles.