Banana Prompt
A fill-in-the-blanks cinematic prompt for your Gemini workflow
- prompt
Banana Prompt is the text-assembly half of the Banana Studio pack. Instead of typing a wall of prose into a generator, you fill labeled sections and it stitches them into a structured prompt for you. It was built for the Gemini image path this pack centers on, but there's nothing Gemini-specific about it - any node that eats a STRING prompt will take its output happily.
Why structured at all? Because the current generation of models reads your prompt through an LLM encoder - Gemini, Z-Image, Flux 2 Klein - and those models genuinely respond to clean field separation. One study after another and a lot of r/comfyui trial-and-error has landed on the same lesson: a well-structured paragraph beats a keyword soup, and labeled blocks are the most reliable way to keep subjects from bleeding together. Banana Prompt gives you that structure with seven labeled slots instead of one giant text box, which is also just easier to edit than a wall of prose.
How it works
Mechanically it's boring in the best way - a pure text assembler, no API, no model, no key needed. medium_or_tech is the one required field. If you fill nothing else, the node returns that string verbatim, so it degrades gracefully into a plain prompt box. Fill any optional section and it formats every non-empty section as Title: body, joined by blank lines. Newlines inside a section get collapsed to spaces, so you can write freeform in a box without the formatting collapsing on you.
The sections that matter
medium_or_tech- the container. Camera, lens, medium, composition. The author's tooltip calls it "camera, lens, medium, composition container," so think of it as the frame you hang everything else in.identify_reference- the interesting one. The tooltip is explicit: "defines the immutable facial identity anchor to ensure consistent character appearance across all generations." That's your character-consistency slot.subject_or_presence- how the subject occupies the frame in this moment: emotion, aesthetic, physical presence.action_or_state- what they're doing.environment- location, time, lighting, the surrounding world.clothing_body- what they wear and how the body is presented.final_style- mood and aesthetic, applied last. The author adds "keep it soft," which is decent advice: let the earlier sections carry the load.
The identify_reference slot is worth a moment, because character consistency is the field's hardest unsolved problem. With Gemini image models, the practical 2026 answer is reference images plus a stable verbal identity - this is the text half of that pairing. Keep the phrasing of identify_reference identical across runs and it acts as the anchor the rest of the prompt hangs off.
Wiring it up
The single output is prompt (STRING). It's meant to feed the prompt input on the pack's BananaStudio node, but it works just as well into a CLIP Text Encode, a Z-Image prompt slot, or any other text input you've got.
The one thing to know
There are no failure modes to troubleshoot here because there's no processing to speak of - the only way to "break" it is to expect magic it doesn't have. It doesn't enhance your prompt, it doesn't expand variables, and it won't fill in sections you left blank (it just omits them). If you want variable substitution, that's the pack's PromptEditor node instead. This one is purely about keeping your prompt writing structured and repeatable - and for a reusable character pipeline, that's exactly what you want.
Install
It ships inside tjcccc/comfyui_banana_studio, so the install is the pack install:
cd ComfyUI/custom_nodes
git clone https://github.com/tjcccc/comfyui_banana_studio.git
Restart ComfyUI (or find "Banana Studio" in ComfyUI Manager). No model files, no extra pip dependencies - the whole pack declares zero runtime deps. It needs ComfyUI ≥ 0.19.3. And unlike the pack's generator node, this one needs no Gemini key at all.
Inputs (7)
| Name | Type | Default | Description |
|---|---|---|---|
| medium_or_tech | STRING | Medium / Tech Camera, lens, medium, composition container. | |
| identify_referenceopt | STRING | Identify Reference Defines the immutable facial identity anchor to ensure consistent character appearance across all generations. | |
| subject_or_presenceopt | STRING | Subject / Presence Describes how the subject emotionally, aesthetically, and physically occupies the frame in this specific moment. | |
| action_or_stateopt | STRING | Action / State What the subject is doing or current state. | |
| environmentopt | STRING | Environment Location, time, lighting, surrounding world. | |
| clothing_bodyopt | STRING | Clothing / Body What the subject wears and how the body is presented. | |
| final_styleopt | STRING | Final Style Mood / aesthetic, applied last. Keep it soft. |
Outputs (1)
| Name | Type | Description |
|---|---|---|
| prompt | STRING | — |