arkennemasis - ChatGPT / Codex / fal AI nodes
Generate images with your ChatGPT subscription - no API key: your Codex CLI login runs gpt-image-2 inside ComfyUI, billed to your ChatGPT plan. Plus 212 fal.ai nodes, one per model, each with a live price badge, a spending cap, reuse of identical paid runs and free result recovery: all Flux (1, 2, Kontext, Flux 3 video), all MiniMax (Hailuo, H3, Speech, Music), all ElevenLabs, all Nano Banana, Seedance 2.5, Kling, Topaz, Qwen, Sync lipsync, SAM 3.1 video segmentation and VEED Fabric - one command adds any new fal model. Also: shared image settings, run-N-of-M shot selection, auto-numbered run folders, caption sidecars, reusable system prompts, and video, TTS and caption utilities. Keywords: ChatGPT, Codex, gpt-image-2, fal, Flux, MiniMax, Hailuo, ElevenLabs, Nano Banana, Seedance, Kling, lipsync, image to video, text to speech.
Nodes (291)
Your cut-out loses its transparency the second you hit Save — this node doesn't
The idea form that starts a fully-automated presenter video
Making the clip exactly as long as the voice-over
Turning one messy LLM answer into a script, a caption, and a filename
Turn the Excalidraw client board into a PNG you can actually look at
Stop guessing your thresholds — measure them
Subtitles that don't look like default subtitle nodes
The review board that shows you the gaps, not just the wins
One chain, N rows, rendered concurrently — the catalogue loop
Checking every reference before you pay for a single image
Writing each image to exactly the filename the CSV promised
The tiny node that keeps a 50-axis graph sane
One chain, N cells, actually concurrent
The cartesian product that knows when to stop
Gpt-image-2 in ComfyUI — with your ChatGPT subscription, no API key
GPT-5 text and vision in ComfyUI, billed to your ChatGPT plan
Check your ChatGPT login before you spend a generation
Every scene's still, one image, no hunting
One accepted image in, a whole delivery ladder out
Fifty-ish voices, eleven languages, five cents per thousand characters
Three references, one sentence, four cents
Qwen Image 3 for text that's actually readable in the picture
Turning a FLUX 3 draft into the real thing
Three cents a second to change a clip you already have
FLUX 3 extend video
Does the join work? 12 cents a second to find out
Two stills in, the shot between them out — FLUX 3 first/last frame
Check the transition for 6 cents a second before you pay for it
Animate one still with FLUX 3 — the photo you already like, now moving
Try three animations of the same photo for the price of one
FLUX 3 with frame-level control
Fix the pacing at 6 cents a second, then render it properly
The good one, and you're paying $0.17–0.29 a second for it
The cheap way to find out if the shot works
14 cents a second vs SeedVR2 on your own GPU
The $1.16-a-second button that finishes a Seedance draft
The frame you approved, now moving with sound
Lock a character across shots with Seedance's reference-to-video
How to hold down the most expensive button in the pack
First frame, last frame, and the model fills the gap
The one that edits and extends video, not just makes it
The simplest node in the pack and the one to start with
A real score for your clip, billed by the rounded-up minute
Identical knobs, different model — so A/B it
The narration node, and the audio tags are the fun part
Half the price, and you probably won't hear it
Keep the performance, swap the voice — for 1.5 cents a minute
Pull the voice out of a noisy clip before anything else
Your clip, another language, same voice — 60 cents a minute
Real word timings for captions, and the $0.22 minimum that bites
Same price as v2, and one knob v2 doesn't have
Foley for a silent clip, at two-tenths of a cent per second
Transcripts you can actually caption with
The cheap transcription node, with a keyterm trick
Two voices, one take, no editing job afterwards
The older, pricier sibling — when it's still the right pick
Eleven widgets and the two that make long narration work
The bulk-rate voiceover node
ElevenLabs Voice Changer in ComfyUI
The fal Flux 2 node in ComfyUI
Instruction editing that keeps image quality, not the face
The half-cent Flux 2 you use to find the prompt that works
Picture-in, picture-out for half a cent a megapixel
The expensive Flux 2 that's actually about the text in your image
Multi-reference editing where you set the shape, not the model
Half a cent per megapixel for the Apache-licensed one
The node you use to preview a fine-tune, not to make pretty pictures
Instructions that actually listen, because guidance isn't baked in
Attach a style LoRA to an edit, for double the price
Text-to-image with your own weights attached
A cent an edit, and eight steps instead of four
The cheapest way to bring your own style to an edit
Styles, characters and consistency, at a cent a megapixel
The daily driver, at six tenths of a cent per megapixel
The checkpoint the Klein LoRA ecosystem is actually trained on
The editor with a working negative prompt
Two cents a megapixel to edit with your own weights
Sample a fine-tune you couldn't run, for two cents
The local editing default without the 20GB
Edit a photo with your own LoRA, on someone else's GPU
Text-to-image with your own weights, $0.015 a megapixel
The 32B quality ceiling with your own style weights
4 reference images, your LoRA, and the best open editor in the room
Product shots on white, in, real scenes out
Furnish an empty room without a van full of sofas
A whole style in one node, and a trigger word you'll mistype
Panel-ready style, and no, it won't restyle your photo
Extend a headshot into a body, and hope the shoes match
The surreal, over-saturated look, with a trigger word you never see
Turn one photo of an object into a turntable, in degrees
The photorealism switch that only goes one way
Orbital imagery on demand, at $0.021 a megapixel
One node, one look, and a surprisingly useful default
Person plus garment in, worn outfit out
The knob-less flagship, and what $0.07 buys you
Multi-reference editing at the top of the price list
The API-only flagship, at $0.03 for the first megapixel
The API-only flagship with reference images attached
Flux 2 Pro Outpaint has no prompt box, and that's the whole point
Eight tenths of a cent per megapixel, no GPU required
Instruction editing at a fraction of a cent, if you mind the aspect ratio
Flux 3 Action SO-101 is a robot arm node, and yes it really is in your ComfyUI menu
Structure control without installing a single ControlNet
Two images in, one composition out
Hold the room, change everything in it
Same room, new materials, one dial apart
The node you use when the batch is bigger than your afternoon
The default is 0.95, and that's not a typo
No prompt box, because Redux doesn't take instructions
LoRAs, ControlNet, IP-Adapter and negatives, all behind one node
A change map instead of a mask, so edits fade instead of stopping
Img2img with the whole extension drawer open
A real mask, on a hosted Flux, with a caveat worth reading twice
Edit a photo by taking it apart and putting it back differently
Two and a half cents for an edit, no 12 GB checkpoint
The node for the LoRAs nobody ported to the newer editors
Mask plus instruction, from one reference
Using an editing model as a generator, on purpose
The aesthetic fix for Flux's plastic look, without the 12GB download
A strength dial that goes all the way to 'remake this'
Stack your Civitai downloads onto a model you never download
Restyle a photo with a LoRA you'll never download
Mask one region, leave the rest of the pixels alone
No prompt box, because the picture already is the prompt
Your Civitai LoRAs, on a GPU that isn't yours
Keep the edges, change everything they contain
Control the layout with depth instead of edges
The inpainter that can be handed elements to insert
Style transfer without training anything
The everyday masked-inpaint node, with LoRAs attached
Change one thing in a photo by describing the change
Double the price for better instructions and legible text
The expensive one — six references, best instruction-following
No reference image, same premium instruction-following
Six image sockets, one instruction that can reference them
Kontext's prompt-following without the reference photo
The closed model with two knobs and a per-megapixel bill
One image, optionally a few words, and the premium stack's opinion
The six-cent image you'll never fit on your own GPU
Your own LoRA, except it lives on someone else's server
Hand it a picture, get the same picture back, but different
Paint over the ex, get a photo that was never at the wedding
The inpainting endpoint for people who don't want to run a 12B model
Masking, but with your own concept baked in
Virtual try-on that doesn't need a model, a studio, or twelve garments
Put a real face in a prompt without wrecking the prompt
Three-tenths of a cent a megapixel, which is basically free
No prompt box, and that's the entire point
The aesthetics-tuned Flux that wants shorter prompts than you're used to
Img2img where the default strength is 0.95, so read it first
Keep one specific thing, generate everything around it
Ten cents a megapixel, and it will invent detail if you let it
Your photo, someone else's dance
Ten cents for six seconds, and 512p is the catch
1080p, fixed at $0.48 a clip, with an end-frame trick
Two inputs, 48 cents, no starting frame to hide behind
768p or 512p, and you choose the price too
27 cents a shot, and the cheapest way to test a scene idea
1080p Motion From One Still for a Flat $0.33
The $0.19 Draft Clip You Should Be Rendering First
$0.49 a Shot, and It's a Hammer, Not a Chisel
$0.49 for a Clip With No Anchor
768p, and the 6-or-10-Second Dial
The Half-Price Way to Find Out If Your Prompt Works
A Penny an Image, and Yes That's the Whole Point
Same Face, New Scene, One Cent a Try
Your Lyrics, Someone Else's Song, 3.5 Cents a Take
3 Cents a Song, and a Box That Lies About Its Name
3 Cents, the Same Price as v1.5, Without the Naming Trap
$0.15 a Song, Five Times the Price of v2 — Worth It?
The Version With the Bigger Tag Vocabulary
10 Cents per 1,000 Characters, Better Than the Name Suggests
6 Cents per 1,000 Characters, Persian and Tagalog Included
$0.10 per 1,000 Characters, No Voice Clone Required
The $0.06 Draft Voice You Keep Coming Back To
Pause Markers, Loudness Normalization, 10,000 Characters a Call
Pause Markers and Normalization for 6 Cents per 1,000 Characters
Where It Starts Sighing and Laughing on Command
Narration at six cents a minute
Two widgets, fifty cents a clip
The node where [Zoom in] is a real instruction
Your frame, MiniMax's camera move
Fifty cents to make your still move
The same node twice, and how to A/B it for 50 cents
Animating a still with the 'live' checkpoint
One photo, the same face in every clip
Ten seconds of audio and $1.50 buys you a voice
Describe a voice you've never recorded, for $3
25 closed-model images for a dollar
The image model that can look things up first
Six images, a video, a PDF and one sentence
Change the thing by describing the change
15 cents an image, and the one to pick for text
Multi-reference compositing at 4K
Segment every frame for a penny per 16 frames
Lip sync that also acts — 15 seconds at a time
The $3-a-minute workhorse
When the mouth is the shot and it has to be right
Great mouth, $8-a-minute mouth
Gemini image gen for about a nickel
The cheap Gemini image node you'll actually spam
Six image sockets and a sentence of instruction
Get that paid render back without paying again
2K with sound, and the resolution box that decides your bill
Consistency, at a JSON box's worth of pain
Turn a Blender grey-box into footage
Orbit a still with a JSON keyframe list
Keep the take going without a visible seam
The one that follows your prompt better
Splice a generated scene into a take you already shot
A photo plus a voice track, for a fraction of Sync's price
Keep the same character across every shot
Video in a style the prompt can't ruin
An animatic that moves like pencil on paper
PS1-era 3D motion in one prompt
Hand-painted animation, no style keywords
Tape damage you can dial, from subtle to unwatchable
Start here, and leave the expansion mode alone until you need it
Hand it one frame, get 15 seconds and a soundtrack
What \"turbo\" buys, and what it costs per second
12 files in, and a prompt that has to learn to count
The LoRA is a JSON field, not a checkpoint
The full-fat endpoint, 5 to 15 seconds at up to 4K
Lock a style in, at 25% above the base rate
Two cents a second for a full song, and lyrics is not a suggestion
The edit model that obeys, and one dropdown that triples the bill
You restarted ComfyUI mid-render. fal Recover Result gets the file back for free.
Keep the performance, swap the voice
Ten cents per ten seconds to not repaint your footage
One photo plus a voice track equals a talking head
VEED Fabric 1.0 Fast costs more per second than the slow one. That's not a typo.
Skip the TTS node, put the script straight on the node
This one is a voice agent, not a transcriber
Two paths to an image — one of them costs nothing
One scene, start to finish, and not an ounce of VRAM more
HTML projects to MP4, and Node.js is the real dependency
One settings node to drive every image generator you own
Teach the model to write its own instructions
The write that makes a run resumable
Resume a run without paying for it twice
Join the film you never finished rendering
A GGUF in your graph, in its own process, with the VRAM handed back
Turning a jittery segmentation mask into something you can composite
Making the generator's output follow the shape of your upload
Land every voice-over on the same mark — no more 1.5s gaps next to 4.4s gaps
The shot length that exactly covers the voice-over
Judging a whole fanned-out run on one annotated board
Standing a cut-out presenter in front of your background, your way
The base photo you'll never regenerate again
The node that picks one prompt from the batch and bolts the locks on
A checksum for the part of the prompt that must never change
The prompt builder that only substitutes
Ask for every variation prompt in one LLM call, then never ask again
The node that actually gets your VRAM back between stages
Teach a vision model to be your QC critic — and make sure it sees both images
Turn the critic's 'looks wrong' into a rewritten prompt
Voice cloning in ComfyUI, no cloud account, no transformers downgrade
The meta-prompt that doesn't know what a product is
Validate the recipe, inject the locks, freeze it — or stop the run
The one human checkpoint in an automated pipeline
Pull one region's mask out of the plate without naming it
Repaint one region to an exact hex and not a pixel outside it moves
A contact sheet your client can actually click, plus the Excalidraw matrix
The barrier that stops the summary from running before the run is over
One numbered output folder per run, wired into every save node
Your Google Sheet state layer, minus the Google Sheet
First-time pass rate
One panel that answers 'what is this run actually going to do?'
One scene of the plan, by index, as plain values — the piece the loop needs
Render 50 scenes from one chain
Pull scene N out of the plan without failing the whole run
The screenshot service that photographs articles, not consent dialogs
Read a client's variation sheet without guessing what it means
Run only 5 of your 24 API calls — the node that never bills a skipped branch
The reference-image cache that stops a batch from drifting silently
From 'a folder of images' to 'the variations are live' — the WooCommerce CSV
The one-form-per-film node your story agent actually wants
The memory that stops your daily video covering yesterday's story
The node that refuses to let a model invent today's story
Remembering a story only once it's actually a video
State the subject once, keep every prompt gender-neutral
The tiny node that makes a LoRA dataset actually trainable
Fail here and it costs nothing. Fail at cell 380 and it costs 380 images.
The QC step that used to live in the operator's head, now a node
Concat, level speech, duck the music, burn the subs
H3 won't do silence — so this node replaces its soundtrack with your narration
One dropdown, two renderers — and the model you didn't pick never even loads
Save the clip, free the models
Free page capture that gives you the words, not just the picture
Captions that light up the word actually being spoken
The reusable system prompt node that works with any LLM
arkennemasis — ComfyUI Nodes
[!TIP]
🚀 Control local ComfyUI from your web AI client
Arkennemasis MCP connects an MCP-capable web AI client to your local ComfyUI through a configured HTTPS tunnel. Your models and GPU stay on your computer.
Set up the optional gateway from the local MCP setup screen:
- Create, Control & Edit: Ask ChatGPT Web to construct entire workflows, wire nodes, tweak parameters, and fix graphs directly on your live canvas.
- Live Browser Sync: Apply revision-checked changes to an explicitly shared ComfyUI tab, with acknowledgments and bounded undo history.
- Selected Node Development: Enable source editing and maintenance for named custom-node folders, with backups and explicit execution controls.
On Windows portable, 3. Create launchers adds one MCP launcher:
run_nvidia_gpu_fast_fp16_accumulation_with_mcp.bat. Userun_nvidia_gpu_fast_fp16_accumulation.batfor the same fast-FP16 mode without MCP. The standard GPU and CPU BATs remain ordinary ComfyUI launchers. Close the active launcher before switching.AI-client availability and usage limits still apply, as do any paid nodes' own credentials and charges. See the current ChatGPT connection requirements.
📖 Full MCP Guide: Arkennemasis MCP Setup & Usage · All Setup Guides
One pack, one menu (arkennemasis), many AI use cases. 76 nodes today:
| Category | | |
|---|---|---|
| Variation | 35 | a client's spreadsheet plus one photo → a verified, consistently-framed product image library |
| Utility | 17 | web capture, masks, compositing, boards, captions, image and text helpers |
| Video | 11 | per-shot generation, dubbing, measured captions, narration-fitted assembly |
| Avatar | 6 | find a story, write it, speak it, and put a presenter in front of it |
| Image Gen · LLM · Audio | 7 | gpt-image-2, GPT-5 text + vision, local TTS |
Providers and use cases keep growing — each module loads independently, so nothing breaks anything else. And if you already pay for ChatGPT, none of the image or text nodes need an API key.
What you can build
| You want | The pack gives you | |---|---| | Live ComfyUI control via a web AI client | Connect an MCP-capable web client to ComfyUI via Arkennemasis MCP. Inspect, edit and explicitly run a shared canvas, with optional access to selected custom-node source. | | A talking-head news video, every morning, by itself | It finds a story, writes it, speaks it, animates a presenter, cuts them out and stands them in front of a screenshot of the article. It remembers what it already covered. | | A narrated film from one brief | Write the idea once. It plans the scenes, renders each one, makes every shot last as long as its voice-over, then joins them with music and subtitles. | | A product photo library from a spreadsheet | One photo of the product plus a list of colours or finishes in, a full set of images out — same object every time, only the named part changing. | | The same picture, many versions | Different hair, different language, different pose. You supply the list of variants; the pack runs them and puts every result on one board to compare. |
You bring the words, the pack brings the machine. Nothing here tells a model what to write — the prompts are yours. These nodes handle the parts that are the same every time: looping, retrying, staying in budget, keeping the shape, timing captions to speech, writing files where you expect them.
It won't quietly spend your money. You can run 3 of 30 branches and the other 27 never execute, so a paid node upstream is never called. Every run gets its own folder and a log.
⭐ Generate images with your ChatGPT subscription — no API key
If you already pay for ChatGPT, you can generate gpt-image-2 images in ComfyUI without
buying any API credit. Install the pack, run codex login once in a terminal, and the
Codex Image Gen node signs in with your existing ChatGPT account. Images bill against
your plan instead of per call.
ChatGPT Plus at $20/month is enough and is what this pack is developed against.
codex login # once, in your own terminal — not inside ComfyUI
There is no API key field and no OAuth flow in ComfyUI — the node reads the Codex CLI's own login. Drop in Codex Login Status to confirm the account and plan before you run anything. Full details in What each path needs.
🎬 Ready-made workflows
example workflows/ ships complete production workflows ready to drag and drop into ComfyUI:
- Property Walkthrough AI — generates a narrated, captioned property walkthrough video from a folder of listing photos.
- Story Creation using MiniMax H3 — plans scenes from a brief, renders each shot with MiniMax H3, dubs narration, times subtitles, and joins into a finished film.
Nodes
| Menu | Node | What it does | Out |
|---|---|---|---|
| arkennemasis/LLM | arkennemasis Codex LLM (ChatGPT login) | GPT-5 text and vision through your codex login — no API key, billed to your ChatGPT plan. Reads images, and splits a long answer into batches so a 50-scene plan does not have to arrive in one reply | STRING, STRING |
| arkennemasis/fal/Image/… | 106 image models — Flux (1, 2 pro/max/flex/flash/turbo/klein, Kontext, LoRA, Fill, Redux, Canny/Depth, PuLID, LoRA galleries, upscaler), Nano Banana (1, 2, Pro, Lite + edits), GPT Image 2 Edit, MiniMax Image 01, Qwen Image 3 | one node per model on fal.ai, with that model's own settings and a live price badge | IMAGE, STRING |
| arkennemasis/fal/Video/… | 62 video models — MiniMax (Hailuo 02 / 2.3 / Fast, Video 01, H3, H3 Max + styles), Flux 3 video (+ drafts, extend, edit, upscale), Seedance 2.5, Kling v3 motion control, Topaz video upscale, SAM 3.1 video segmentation | video, most with native audio | VIDEO, STRING |
| arkennemasis/fal/Lip Sync/… | Sync Lipsync 2 / 2 Pro / 3 / React-1, MiniMax H3 Max Lip Sync, VEED Fabric 1.0 (+ Fast, + Text) | talking video from a picture or a clip plus audio | VIDEO, STRING |
| arkennemasis/fal/Audio/… | 36 audio models — ElevenLabs (TTS v2.5/v3/v4, dialogue, music, sound effects, speech-to-text, alignment, voice changer, isolation, dubbing), MiniMax (Speech 02–2.8, voice clone/design, Music), Chatterbox, Qwen Audio 3 TTS, xAI Grok Voice | speech, music, sound, transcripts | AUDIO, STRING |
| arkennemasis/fal/Tools | fal · History (free) · Recover Result (free) | every fal run, loaded back for free; a finished request collected by id without paying again | STRING, IMAGE, VIDEO |
| arkennemasis/Image Gen | arkennemasis Image Gen Settings (shared) | one node driving aspect_ratio / quality / run_mode / background / output_format / moderation / timeout_seconds / api_token on many Image Gen nodes at once | ARK_IMAGE_SETTINGS |
| arkennemasis/Image Gen | arkennemasis Codex Image Gen (ChatGPT login) | gpt-image-2 through your codex login — no API key, billed to your ChatGPT plan | IMAGE, STRING |
| arkennemasis/Utility | arkennemasis Codex Login Status | which ChatGPT account this machine will use, and when its token expires | STRING |
| arkennemasis/Utility | arkennemasis System Instructions | reusable system prompt for any LLM node | STRING |
| arkennemasis/Utility | arkennemasis Shot Selector (run N of M) | run only N of M expensive branches — the first N in order, or a random sample from a seed. Unselected branches never execute, so a paid API node upstream is never called | IMAGE |
| arkennemasis/Utility | arkennemasis Subject Line (gender + notes) | one Subject: … line from a gender choice plus free-text notes, wired into every prompt — so the prompts themselves stay gender-neutral and the subject is stated once | STRING |
| arkennemasis/Utility | arkennemasis Text File Save (caption sidecar) | writes <folder>/<filename>.txt next to a saved image — the image/caption pairing training toolkits expect | STRING |
| arkennemasis/Utility | arkennemasis Run Folder (auto-numbered) | <parent_dir>/<folder_name>_001, _002, … — one fresh output folder per run | STRING, INT |
| arkennemasis/Utility | arkennemasis Story Brief / Run Log / Contact Sheet | the brief form, a JSON run log that needs no spreadsheet, and every still of a run on one sheet | STRING, IMAGE |
| arkennemasis/Video | arkennemasis Scene List (the loop) | fans a scene plan out so one chain runs once per scene — 5 or 50, same canvas | lists |
| arkennemasis/Video | arkennemasis Hailuo Scene | one scene start to finish: condition → sample → decode video and audio → mux → save → free | VIDEO |
| arkennemasis/Audio | arkennemasis Qwen3-TTS (voice clone) | local Qwen3-TTS. Text in, speech out; give it 5–30 s of someone speaking and it clones that voice. Runs in a subprocess — see below | AUDIO, STRING |
| arkennemasis/Video | arkennemasis Video Dub (narration over a clip) | swaps a clip's own soundtrack for a narration track, per clip. MiniMax H3 always generates audio and cannot be asked for silence, so the voice has to replace it | VIDEO, STRING |
| arkennemasis/Video | arkennemasis Narration Length (fit the shot to the voice) | measures a rendered narration and returns the shot length that covers it, snapped to H3's frame grid. Wire between the TTS node and the scene node and every shot outlasts its own voice-over | INT, FLOAT, STRING |
| arkennemasis/Video | arkennemasis Narration Fit (stretch or compress speech) | stretches or compresses rendered voice-over via pitch-preserving ffmpeg atempo to land exactly on target_seconds | AUDIO, STRING |
| arkennemasis/Video | arkennemasis Caption Style (font + subtitle style) | one of five subtitle styles, any installed font, colours, outline, box, size, 3×3 position — and an on/off switch | ARK_CAPTION_STYLE |
| arkennemasis/Video | arkennemasis Video Assemble (clips + music + subs) | joins every clip, levels each one's speech, ducks a music bed, burns the captions | STRING, VIDEO |
| arkennemasis/Video | arkennemasis Load Clips (finished clips from disk) | reads a run's finished clips back as a VIDEO list — join a film whose render was interrupted, without re-rendering | VIDEO list |
| arkennemasis/Video | arkennemasis Scene Split / Scene At | pull one scene out of a JSON plan — by parsing it, or by index | STRING |
| arkennemasis/Video | arkennemasis Word Timings | transcribes the finished narration and returns when each word is actually spoken. Moving caption styles need real times; guessing from word count drifts | STRING |
| arkennemasis/Video | arkennemasis Video Model / Video Save | pick one video model without both loading, and write a clip then free the weights | MODEL, VIDEO |
| arkennemasis/Utility | arkennemasis Web Shot (page → picture, text, links) | drives a local Chrome and gives you a screenshot, the readable text, and the links. Free, and it can scroll to a selector first | IMAGE, STRING |
| arkennemasis/Utility | arkennemasis ScreenshotOne | the hosted version of the same job, with ads and cookie banners blocked. Some sites photograph as a consent dialog through a plain browser and correctly through this | IMAGE, STRING |
| arkennemasis/Utility | arkennemasis Mask Refine | turns a per-frame segmentation mask into one you can composite through — a segmenter is right frame by frame and jittery as a video | MASK |
| arkennemasis/Utility | arkennemasis Overlay Subject | stands a cut-out subject on a background at a chosen size and position. The part a hosted avatar service does for you and never lets you adjust | IMAGE |
| arkennemasis/Utility | arkennemasis Match Aspect | measures the image being edited and pins the generator's output to that shape. Without it an "auto" setting says nothing about shape and the model reframes your picture | ARK_IMAGE_SETTINGS |
| arkennemasis/Utility | arkennemasis Option Board | collects a fanned-out run onto one labelled .excalidraw board plus a preview — how you judge a set rather than a picture | STRING, IMAGE |
| arkennemasis/Utility | arkennemasis Purge VRAM | unloads models between stages, so a big image model and a big video model can share one canvas | passthrough |
arkennemasis/Avatar — a story in, a presenter telling it out
Six nodes that turn "cover today's news about X" into a captioned vertical video, with no one watching. They compose with the Video and Utility nodes above rather than replacing them.
| Node | What it does | |---|---| | Avatar Brief | one labelled box per thing you actually decide: the beat, the sources, the voice, the length | | Story Pick | reads the model's choice back out and checks it — a real article, from the list it was shown, not one it invented | | Story History | what has already been covered, so tomorrow picks something else | | Story Record | writes today's story into that history — only once a video file exists. A run that produced nothing covered nothing | | Avatar Script | splits one answer into the spoken script, the caption and the headline | | Avatar Frames | how long the clip must be, measured from the voice that will play over it |
arkennemasis/Variation — the product-variation pipeline
A client's variation spreadsheet plus one locked base photograph in; a verified, consistently framed image library out. Any product, any number of variation axes. The guarantee is that every delivered image shows the same physical object, differing only in the specified attribute.
- 35 modular pipeline nodes: Covers intake (
Sheet Probe,Variation Intake), references (Spec Library), geometry (Plate Lock,Region Mask), recipe compilation, substitution prompt generation, CIELAB recoloring, automated ΔE2000 quality verification, and store export. - Two execution paths: The full ten-stage pipeline for unformatted client sheets, and a shorter CSV catalogue path when prompts already exist.
👉 Full architecture & setup guide: See setup_guides/06_product_variation_pipeline_guide.md.
arkennemasis/Audio — Qwen3-TTS Voice Cloning
Local, offline voice cloning. Input text and a 5–20s voice reference sample to clone any voice. Runs in an isolated child subprocess (vendor/tts_env) with pinned dependencies, completely preventing conflicts with ComfyUI's main Python packages.
👉 Full setup & installation guide: See setup_guides/05_local_voice_cloning_qwen3_tts_guide.md.
Captions
Caption Style feeds Video Assemble. Five styles:
| Style | On screen |
|---|---|
| classic | the whole line at once |
| karaoke | the fill sweeps across the line as it is spoken |
| highlight | the spoken word changes colour |
| underline | the spoken word is underlined |
| word_by_word | one word at a time, nothing else |
Everything but classic marks individual words, so it needs to know when each word is
spoken. A video model gives no word timestamps, so they are estimated from the script
and the clip's real duration, weighted by word length and by trailing punctuation. That
tracks speech closely; it is not frame-accurate, and it drifts if the model ad-libs.
Fonts come from fonts/ (bundled, listed first) and from the machine's own
installed fonts. See that folder's README for what ships and how to add more.
Toggle enabled off on the node for a video with no subtitles at all — every other
setting stays put.
Install
cd ComfyUI/custom_nodes
git clone https://github.com/Hishamahmer/comfyui-arkennemasis
Install the dependencies, then restart ComfyUI:
# portable build:
python_embeded\python.exe -m pip install -r ComfyUI\custom_nodes\comfyui-arkennemasis\requirements.txt
# normal install:
pip install httpx
Or in ComfyUI-Manager → Install via Git URL → paste the repo URL (deps auto-install).
What each path needs
Two ways to reach gpt-image-2, plus fal for everything else. Use whichever you already
pay for.
| Path | Nodes | What it requires |
|---|---|---|
| fal.ai | every arkennemasis fal · node | A fal.ai account and an API key (FAL_KEY). Pay-as-you-go per image / second / minute — each node shows its price. |
| Codex / ChatGPT | Codex Image Gen, Codex Login Status | A paid ChatGPT subscription and the Codex CLI already logged in on this machine. No API key. Images bill against your ChatGPT plan instead of per call. |
Codex path — read this before you try it
- A paid ChatGPT plan is required. The free tier cannot call the hosted image tool. ChatGPT Plus at $20/month is the recommended plan and is what this pack is developed against. Business/Pro plans work too.
- Install the Codex CLI and sign in from your own terminal, on the same machine and
the same user account that runs ComfyUI:
This writescodex login~/.codex/auth.json. The nodes read that file directly — there is no OAuth flow inside ComfyUI and nowhere to paste a password. If you have not runcodex login, the Codex nodes will tell you so and stop. - Check it worked by dropping in the Codex Login Status node — it reports the signed-in account, the plan and when the token expires.
Availability is account-dependent: not every ChatGPT plan or region can call the hosted image tool. The node says so plainly rather than failing cryptically.
Running several ChatGPT accounts? Give each its own CODEX_HOME folder and set the node's
codex_home per node. The Codex Image Gen node's account output names the signed-in
email, so you can see which login produced an image.
fal.ai nodes
One node per model. Each fal model is its own node with exactly that model's inputs,
built from the model's file in fal_provider/models/. Under
arkennemasis/fal/: Image, Video, Lip Sync and Tools.
The key. Put one line in the .env file of your ComfyUI install — in the portable build
that is the folder holding run_nvidia_gpu.bat:
FAL_KEY=your-key-from-fal.ai/dashboard/keys
The nodes read it from there and nowhere else — there is no key box on the nodes, so a key can never end up inside a workflow file or an image you share. Get a key at https://fal.ai/dashboard/keys.
Money.
- Each node carries a live price badge that follows its settings (resolution, duration, number of images, quality…), and its title shows the model's rate.
max_cost_usd(default $20) is a cap: when the estimate for a run is above it, the node stops before anything is uploaded or sent.0removes the cap.- The estimates are tested against fal's own prices and worked examples, and the pre-run check measures the real clips and pictures you connect.
max_concurrent(default 1) is how many paid fal calls may run at the same time, across every fal node in a run. At 1 they go one after another; raise it (up to 32) to run several side by side. If one fal node fails, the ones still waiting are not started.- While it runs, the node shows its status (queued / running / seconds / estimate). ComfyUI's Cancel also cancels the request on fal.
- Every request is written to
output/fal/_requests.jsonl. If ComfyUI is restarted while a job is running on fal, fal Recover Result collects it by its id — for free. - Results are saved to
output/fal/<model>/and shown on the node.
Adding a model is one command (free — it reads fal's public pages), then restart ComfyUI:
python_embeded\python.exe ComfyUI\custom_nodes\comfyui-arkennemasis\fal_provider\add_model.py https://fal.ai/models/<owner>/<model>
See fal_provider/README.md for how the model files and the
price rules work.
Advanced Utilities & Execution Controls
Arkennemasis provides dedicated utility nodes to manage execution flow, prevent runaway costs, and organize outputs:
- Shared Image Settings (
Image Gen Settings): Drives aspect ratio, quality, moderation, and timeouts across multiple generator nodes simultaneously. - Rate Limits & Concurrency (
run_mode): Switches between sequential execution (one at a time) to avoid 429 rate limit bans, and parallel execution (all at once,max_concurrent). The fal nodes have their ownmax_concurrent(default 1: one paid fal call at a time across the whole run). - Automated Backoff Retries: Automatically retries 429 rate limits, 5xx server drops, and interrupted network streams with exponential backoff.
- Lazy Branch Gating (
Shot Selector): Evaluates only the first N branches; unselected branches are never evaluated and never billed. - Auto-Numbered Output Folders (
Run Folder): Generates dynamically incremented run folders (run_001,run_002) so outputs stay grouped. - Caption Sidecars (
Text File Save): Writes<filename>.txtcaption pairing files alongside generated images.
👉 Full technical guide: See setup_guides/07_advanced_utilities_and_concurrency.md.
Notes
- fal calls are paid — each run bills your fal account (see the node's price badge).
- Long runs poll (no fixed timeout) and run off the UI thread, so ComfyUI stays responsive and Cancel works. A spinner + elapsed-time badge shows on the node while it runs.
Example workflows
Canvas-format workflow exports live in example workflows/. See the
README there for conventions — most importantly never commit an api_token, since
workflow JSON stores widget values verbatim.
Developer & Contributing Guide
For the complete repository file tree, menu category architecture, adding new providers in 3 steps, reusable helpers, and developer safety rules:
👉 Developer Guide: See setup_guides/08_developer_and_contributing_guide.md.
License
MIT