Nodes/Phi-3-mini in ComfyUI/🏖️Phi3mini 4k
ComfyUI Node

🏖️Phi3mini 4k

A local LLM prompt writer inside ComfyUI — no API key, no cloud, no bill

By ZHO-ZHO-ZHO·Created 2 years ago·Updated 2 years ago· 207
🏖️Phi3mini 4k
  • model
  • tokenizer
  • text
promptWhat is the meaning of life?
system_instructionYou are creating a prompt for Stable Diffusion to generate an image. First step: understand the input and generate a text prompt for the input. Second step: only respond in English with the prompt itself in phrase, but embellish it as needed but keep it under 200 tokens.
temperature0.00

Despite the name, this node doesn't make images. It writes text - specifically, it turns a rough idea into a polished English prompt you can hand to CLIPTextEncode. That's why it keeps appearing in workflows next to SDXL checkpoints: it's LLM-assisted prompting as a node, the practice that went from "open a browser tab with ChatGPT" to "drag a node into the graph" (the KB tracks that shift - dedicated local prompt-writing nodes are now a whole category). No API key, no cloud round-trip, no per-token bill. That's the entire appeal.

It wraps Microsoft's Phi-3-mini-4k-instruct, a 3.8-billion-parameter model under an MIT license. The author's launch pitch - "comparable to GPT-3.5 and Mixtral 8x7B" - is 2024 marketing; it'll give you solid prompt expansions, not a brilliant novelist. Treat it as a handy local helper, which is all you need it to be.

How it works

The node builds a chat-style message list - system instruction plus your prompt - and runs a transformers text-generation pipeline. Generation is capped at 500 new tokens and return_full_text is false, so you get back only the freshly written text, not a transcript. Chinese in, English prompt out (the author's headline feature) is just the default instruction doing its job.

The inputs that matter

Five inputs, and only a few you'll actually touch:

  • model / tokenizer - from the 🏖️Phi3mini 4k ModelLoader node. No loader, no text.
  • prompt (multiline) - your rough idea. "girl in a rainstorm, cinematic lighting" level of input is fine; the model does the embellishing.
  • system_instruction (multiline) - the default is a working Stable Diffusion prompt-writer instruction. This is where you customize the job: captioning an image, translating, forcing a tag style for anime checkpoints, whatever.
  • temperature (0–1, default 0) - and here's where people get burned. The node hardcodes do_sample=False, and greedy decoding ignores temperature entirely. The slider is decoration. Don't waste turns twiddling it.

The output and where it goes

One output: text (STRING). Wire it into any text display node - the built-in Show Text, or DisplayText_Zho if you have the author's ComfyUI-Gemini pack. In the bundled CosXL workflow it feeds straight into CLIPTextEncode, so the whole loop is: loader → this node → text encoder → sampler.

A note on which checkpoints to pair it with: it's an LLM writing natural language, and natural language is exactly what LLM-encoded models want. For SDXL-era checkpoints it adds polish and variety to your prompts; for models like Anima or Z-Image it's arguably a more natural fit than tag lists.

Installing and first run

Install with the pack. ComfyUI Manager → search "Phi-3-mini", or:

cd ComfyUI/custom_nodes
git clone https://github.com/ZHO-ZHO-ZHO/ComfyUI-Phi-3-mini
cd ComfyUI-Phi-3-mini
pip install -r requirements.txt   # transformers>=4.40.0

then restart ComfyUI. The one genuinely heavy step is the model: the loader auto-downloads it from Hugging Face on first run (several gigabytes, CUDA required - the loader hardcodes device_map="cuda"). The README's upgrade path if transformers is too old: pip uninstall -y transformers && pip install git+https://github.com/huggingface/transformers.

One more quirk: prompt is a text box on the node, not a wire input - you type into it. The whole point of the node is that you hand it a sentence and it hands back a better one; if you want to keep the conversation visible across runs, use the Chat variant instead.

Category🏖️Phi3mini

Inputs (5)

NameTypeDefaultDescription
modelPHI3
tokenizerTK
promptSTRINGWhat is the meaning of life?
system_instructionSTRINGYou are creating a prompt for Stable Diffusion to generate an image. First step: understand the input and generate a text prompt for the input. Second step: only respond in English with the prompt itself in phrase, but embellish it as needed but keep it under 200 tokens.
temperatureFLOAT0.000–1

Outputs (1)

NameTypeDescription
textSTRING