Extensions/ComfyUI Flamin Galah
ComfyUI Extension

ComfyUI Flamin Galah

ComfyUI nodes that turn filmmaker-style fields and images into MiniMax H3 T2VA/I2VA prompts via local Ollama or the xAI Grok API.

By flamin-galah·Created 12 days ago·Updated 2 days ago· 0
flamin-galah/ComfyUI-Flamin-Galah
Nodes—
On cloudLocal install
Stars0
Updated2 days ago
Readme
<p align="center"> <img src="docs/flamin-galah-logo.png" alt="Flamin Galah logo" width="250"> </p> <h1 align="center">ComfyUI Flamin Galah</h1> <p align="center"> Custom ComfyUI nodes that turn filmmaker-style fields (and optional images) into official <strong>MiniMax H3</strong> T2VA / I2VA prompts. </p> <p align="center"> <img src="https://img.shields.io/badge/ComfyUI-custom%20nodes-orange" alt="ComfyUI"> <img src="https://img.shields.io/badge/Ollama-local-green" alt="Ollama"> <img src="https://img.shields.io/badge/xAI-Grok%20API-black" alt="Grok"> <img src="https://img.shields.io/badge/license-MIT-blue" alt="MIT"> </p>

What this pack does

MiniMax H3 wants a rigid three-field prompt, not a free-form idea dump:

integrated_multimodal_description: [Shot 1] ...
overall_soundscape: ...
non_diegetic_music: ...

For image-to-video (I2VA) it also wants a first-frame lock line:

For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced.

These nodes fill that structure for you.

| Node | Category | What it produces | | --- | --- | --- | | Flamin Galah NSFW Prompt Generator | prompt/Flamin Galah | Structured fields → local Ollama → H3 T2VA prompt. Connect an image to switch to I2VA. | | Flamin Galah Image Prompter | prompt/Flamin Galah | Image → Ollama vision + writer models → H3 I2VA block | | Flamin Galah Grok API Prompter | prompt/Flamin Galah | Image → xAI Grok API → H3 I2VA block |

<p align="center"><em>Flamin Galah NSFW Prompt Generator — structured action, camera, style, lighting, dialogue, sound, and Ollama model.</em></p> <p align="center"> <img src="docs/nsfw-prompt-generator.png" alt="Flamin Galah NSFW Prompt Generator node in ComfyUI"> </p> <p align="center"><em>Flamin Galah Image Prompter — local Ollama vision + writer models to an H3 I2VA block.</em></p> <p align="center"> <img src="docs/image-prompter.png" alt="Flamin Galah Image Prompter node in ComfyUI"> </p> <p align="center"><em>Flamin Galah Grok API Prompter — same I2VA shape via the xAI Grok API.</em></p> <p align="center"> <img src="docs/grok-api-prompter.png" alt="Flamin Galah Grok API Prompter node in ComfyUI"> </p>

Install

  1. Copy the ComfyUI-Flamin-Galah folder into ComfyUI/custom_nodes/.
  2. Fully restart ComfyUI (a simple refresh is not enough the first time).
  3. Add nodes from Add Node → prompt → Flamin Galah.
  4. If you already had an older Flamin Galah node on the canvas, delete it and add a fresh one so widgets match the current schema.

No extra Python packages beyond a normal ComfyUI install (torch, numpy, Pillow).


Setup: local Ollama (prompt generator + image prompter)

  1. Install and start Ollama so it listens on http://localhost:11434.

  2. Pull a writer model. Uncensored / adult-capable models work best for NSFW scenes:

    ollama run jimscard/adult-film-screenwriter-nsfw
    

    Also reported as working well:

    ollama pull fluffy/l3-8b-stheno-v3.2
    
  3. For I2VA / the Image Prompter node, also pull a vision model:

    ollama pull llava
    ollama pull qwen2.5-vl
    ollama pull llama3.2-vision
    

The ollama_model / vision_model / prompt_model dropdowns are filled from GET /api/tags when the node class loads. Start Ollama before launching ComfyUI, or refresh / restart after pulling models.


Setup: Grok API Prompter

The Grok node talks to https://api.x.ai/v1/chat/completions.

Provide a key in one of these ways:

  • paste it into the node's api_key widget, or
  • set XAI_API_KEY or GROK_API_KEY in the environment that launches ComfyUI.

Selectable models: grok-4.7, grok-4.6, grok-4.5, grok-4, grok-2-vision-1212.


Node reference

1. Flamin Galah NSFW Prompt Generator

Output: h3_prompt (STRING)

Connect that string into whatever MiniMax H3 / text node you use.

Required widgets

| Widget | Role | | --- | --- | | action | Scene / subject / motion. Tall multiline field (frontend forces ~180px). | | camera | Presets: static wide, POV, slow push-in, tracking, dolly zoom, handheld, low-angle, aerial, orbit, close-up, pan, intimate close-up, slow tilt up. | | style | live-action cinematic, photorealistic, erotic film, softcore cinematic, glamour, anime, cyberpunk, film noir, documentary. | | lighting | Practical / mood presets (soft red, candlelight, neon wet streets, golden hour, moonlight, chiaroscuro, …). | | duration_seconds | 4–15. Written into the shot as “The shot lasts about N seconds.” | | include_dialogue | If on, dialogue is copied verbatim into <d>[Language] exact words</d>. Nothing is invented. | | dialogue_speaker / dialogue / dialogue_language | Speaker label, exact line, language tag. | | soundscape | Diegetic audio only (breathing, fabric, rain, room tone). Never merged into the visual field. | | music | Non-diegetic score. Use N/A for silence. | | extra_details | Texture, DoF, atmosphere. | | ollama_model | Live list from your Ollama instance. |

Optional

| Widget | Role | | --- | --- | | image | Connecting an IMAGE switches the node to I2VA: vision model describes the frame, then the writer evolves it with your action fields and prepends the official first-frame line. | | vision_model | Ollama vision model used only when image is connected. | | ollama_url | Default http://localhost:11434. | | temperature | Default 0.7. | | negative_notes | Guidance for the LLM only — not written into integrated_multimodal_description. | | lora_tags | Optional trigger words appended for LoRA-aware pipelines. |

Rules the writer is instructed to follow

  • English except spoken dialogue and on-screen text.
  • First shot starts with [Shot 1] (no timestamp).
  • Camera movement is explicit.
  • On-screen text in English double quotes.
  • No sound inside integrated_multimodal_description.
  • No “Avoid / negative” lines in the visual field.

2. Flamin Galah Image Prompter (Ollama)

Local-only two-model pipeline:

  1. Vision model describes the connected image (identity, clothing, pose, lighting, set — no invented action, no sound).
  2. Prompt model turns that description + extra_description into a continuous first-frame-forward shot paragraph.
  3. A third short call writes overall_soundscape from the shot (ambient + physical + non-verbal human sound only).

Inputs

| Name | Type | Notes | | --- | --- | --- | | image | IMAGE | First frame. Batch item 0 is used. | | extra_description | string | Optional “what happens next” / camera move. | | vision_model / prompt_model | Ollama dropdowns | Vision vs writer. | | temperature | float | Default 0.2 (more conservative than the generator). | | lora_tag | string | Optional. | | ollama_url | string | Default localhost. |

Output: image_description — full I2VA block including the first-frame line, [Shot 1] Live-action, cinematic, …, soundscape, and an empty non_diegetic_music: for you to fill.


3. Flamin Galah Grok API Prompter

Same I2VA shape as the Ollama Image Prompter, but every call goes to xAI.

Inputs: image, extra_description, api_key, grok_url, grok_model, temperature (default 0.2).

Flow: vision describe → enhance with your extra prompt → soundscape from the still. Identity, clothing, body, set, and lighting from the photo are preserved; your extra text becomes the continuation after frame 0.


Example T2VA output

integrated_multimodal_description: [Shot 1] Photorealistic, high-contrast dramatic light. A slim woman with huge firm breasts sits on an office chair, legs open wide, looking into lens. The shot lasts about 5 seconds. Camera holds a static wide shot. Subtle skin texture, shallow depth of field.

overall_soundscape: Soft breathing, fabric sliding on the chair, distant rain against the window.

non_diegetic_music: N/A

Example I2VA header

For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced.

integrated_multimodal_description: [Shot 1] Live-action, cinematic, ...
overall_soundscape: ...
non_diegetic_music:

Typical graphs

Text-to-video

Flamin Galah NSFW Prompt Generator → MiniMax H3 T2VA / prompt consumer

Image-to-video (local)

Load Image → Flamin Galah Image Prompter → H3 I2VA
or connect the same image into the generator’s optional image input.

Image-to-video (cloud)

Load Image → Flamin Galah Grok API Prompter → H3 I2VA


Troubleshooting

| Symptom | Fix | | --- | --- | | Dropdown says “Ollama not running” | Start Ollama, then restart ComfyUI so _get_ollama_models() can hit /api/tags. | | Could not reach Ollama | Check ollama_url, firewall, and that the daemon is up. | | Empty / weak I2VA description | Use a real vision model (llava, qwen2.5-vl, llama3.2-vision), not a text-only tag. | | Grok node errors on key | Paste api_key or export XAI_API_KEY / GROK_API_KEY for the ComfyUI process. | | Old widgets / missing fields | Delete the node from the graph and add a new one after updating the folder. | | Action box too small | The pack ships web/flamin_galah.js, which forces a taller action widget and a minimum node size. |


Repo layout

ComfyUI-Flamin-Galah/
├── __init__.py                  # NODE mappings + WEB_DIRECTORY
├── LICENSE                      # MIT
├── README.md
├── docs/
│   ├── nsfw-prompt-generator.png
│   ├── image-prompter.png
│   └── grok-api-prompter.png
├── nodes/
│   ├── __init__.py
│   ├── node_nsfw_prompt_generator.py
│   ├── node_image_describer.py
│   └── node_grok_api_image_describer.py
└── web/
    └── flamin_galah.js          # larger action textarea on the generator

License

MIT © 2026 flamin-galah