MiniMax H3 Prompt Workbench
An agent writes your prompt in a box, then waits for permission
- video_settings
- resources
- context_latent
- vae
- prompt
Why this node exists
MiniMax H3 is a 33B omni-modal model that takes text, image, video and audio as one context and generates 4–15s clips at up to 2K/24fps with native stereo audio - dialogue and room tone made with the picture, not bolted on. Prompting that means writing a shot list, and the blank page is the hard part.
The Prompt Workbench is a text box with an optional writing partner attached. It has one output, prompt (STRING), which you wire into whatever you'd otherwise type into your H3 text encode - and an output node, so it runs on Queue with nothing downstream.
What it actually does
The interesting part is what a normal run does not do. video_settings, resources, context_latent and vae are all declared lazy, and the node's lazy-status check returns nothing unless a preparation job is in flight - so a plain Queue never evaluates the bundles or the VAE. It echoes finalized_prompt and stops. Drafting is button-driven, not queue-driven.
Hit Generate prompt and the frontend asks the backend for a job; the backend gathers the H3 settings, resource bundle and motion context it needs, and a CLI agent runs inside a Docker image built on python:3.13-slim-trixie with Bubblewrap. That agent gets exactly three tools - get_context, read_image, read_skill. No shell, no edits, no web, no subagents, no Docker socket, no GPU, no mount of your ComfyUI install. It's the opposite bet to the usual local-LLM enhancer: no abliterated model on your card, no regex scrubbing chat preamble. The failure you're protected from isn't a leaked "Here is your enhanced prompt:", it's the agent doing something.
The draft appears in Generation results → Output prompt, where you edit it and Apply to Workbench, or use Apply output on the node; applying replaces finalized_prompt as one undoable edit, while closing the dialog changes nothing. Generation never queues your sampler or save nodes - it prepares a prompt, not a render.
Inputs worth touching
- agent -
codex(default) orgrok, set up under Settings → Arisu Nodes → Prompt Workbench → Agents: build the image, do the device-code login, pick a model and reasoning effort (only low, medium and high are offered). - skill -
bundled:with-refby default (Ref2VA / Hybrid: you have keyframes or references),bundled:no-reffor T2VA / I2VA / FL2VA / L2VA. Drop a folder with aSKILL.mdatuser/__arisu_nodes/skills/<skill-name>/and it shows up ascustom:<skill-name>- your house style, on demand. - requirements - the real brief, multiline. trigger_words - text you want present verbatim. motion_notes - what the previous clip did.
- context_length - 5, 22, 39 or 56 frames, default 22; at 24fps that's about a second of previous clip. audio_context_length - 0–240, descriptive only;
0follows the video length and no audio latent is decoded here. - Optional:
video_settingsandresourcesfrom the other Arisu H3 nodes, pluscontext_latentandvae- both required to enable Motion Context.prepare_jobis the hidden job ticket; don't touch it.
Install
ComfyUI Manager → search ComfyUI-Arisu-Nodes → install → restart. Manual:
cd ComfyUI/custom_nodes
git clone https://github.com/swqa7697/ComfyUI-Arisu-Nodes.git
Needs ComfyUI ≥ 0.30.0 and Python ≥ 3.10; the pack is written against ComfyUI's V3 node API (io.ComfyNode / io.Schema). It declares two real runtime dependencies, Pillow and PyAV - check they landed in ComfyUI's interpreter, not your system Python. First startup writes user/__arisu_nodes/config.arisu.jsonc and a skills directory; the nodes live under Add Node → Arisu Nodes.
Agent generation additionally needs Docker with Linux container support that the ComfyUI process can reach, plus a policy permitting the CLI sandbox's nested user namespaces. No Docker, no agent: the left column disables itself and finalized_prompt stays editable, which beats a half-broken node.
Where people get burned
- Expecting Queue to draft a prompt, or Generate to render a clip. Neither. Queue emits the text that's already there, and generation never queues downstream nodes.
- Losing a draft. A new generation clears the previous output immediately, and activity plus unapplied output live outside node data - gone at shutdown. Apply it.
- Motion Context greyed out. Both
context_latentandvaeare required, and arbitrary latent producers are rejected: a stock latent fails withexpected H3 Motion Context Load Latent output. - Docker installed, generation still dead. It has to be the daemon the ComfyUI process can reach, with user namespaces allowed; when the sandbox can't be enforced, generation fails visibly rather than falling back to an unrestricted agent. Provider accounts are shared by everyone on that instance; credentials live in Docker volumes, not workflows.
- The legal bit nobody puts in tutorials. The H3 weights are geofenced - the Community License excludes the US, EU, UK and South Korea. It's also 33B with ~42.5GB of reported full weights and no verified consumer floor.
Inputs (15)
| Name | Type | Default | Description |
|---|---|---|---|
| agent | COMBO | codex | 2 options: codex, grok |
| skill | STRING | bundled:with-ref | — |
| context_length | COMBO | 22 | 4 options: 22, 5, 39, 56 |
| audio_context_length | INT | 240–240 | — |
| motion_notes | STRING | — | |
| reference_notes | STRING | {} | — |
| trigger_words | STRING | — | |
| requirements | STRING | — | |
| finalized_prompt | STRING | — | |
| video_settingsopt | ARISU_MINIMAX_H3_VIDEO_SETTINGS | — | |
| resourcesopt | ARISU_MINIMAX_H3_RESOURCES | — | |
| context_latentopt | LATENT | — | |
| vaeopt | VAE | — | |
| prepare_jobopt | STRING | — | |
| motion_enabledopt | BOOLEAN | true | — |
Outputs (1)
| Name | Type | Description |
|---|---|---|
| prompt | STRING | — |