Nodes/ComfyUI-Arisu-Nodes/MiniMax H3 Prompt Workbench
ComfyUI Node

MiniMax H3 Prompt Workbench

An agent writes your prompt in a box, then waits for permission

By swqa7697·Created 25 days ago·Updated 5 days ago· 8
MiniMax H3 Prompt Workbench
  • video_settings
  • resources
  • context_latent
  • vae
  • prompt
◄agentcodex►
◄skillbundled:with-ref►
◄context_length22►
◄audio_context_length24►
◄motion_notes►
◄reference_notes{}►
◄trigger_words►
◄requirements►
◄finalized_prompt►
◄prepare_job►
◄motion_enabledtrue►

Why this node exists

MiniMax H3 is a 33B omni-modal model that takes text, image, video and audio as one context and generates 4–15s clips at up to 2K/24fps with native stereo audio - dialogue and room tone made with the picture, not bolted on. Prompting that means writing a shot list, and the blank page is the hard part.

The Prompt Workbench is a text box with an optional writing partner attached. It has one output, prompt (STRING), which you wire into whatever you'd otherwise type into your H3 text encode - and an output node, so it runs on Queue with nothing downstream.

What it actually does

The interesting part is what a normal run does not do. video_settings, resources, context_latent and vae are all declared lazy, and the node's lazy-status check returns nothing unless a preparation job is in flight - so a plain Queue never evaluates the bundles or the VAE. It echoes finalized_prompt and stops. Drafting is button-driven, not queue-driven.

Hit Generate prompt and the frontend asks the backend for a job; the backend gathers the H3 settings, resource bundle and motion context it needs, and a CLI agent runs inside a Docker image built on python:3.13-slim-trixie with Bubblewrap. That agent gets exactly three tools - get_context, read_image, read_skill. No shell, no edits, no web, no subagents, no Docker socket, no GPU, no mount of your ComfyUI install. It's the opposite bet to the usual local-LLM enhancer: no abliterated model on your card, no regex scrubbing chat preamble. The failure you're protected from isn't a leaked "Here is your enhanced prompt:", it's the agent doing something.

The draft appears in Generation results → Output prompt, where you edit it and Apply to Workbench, or use Apply output on the node; applying replaces finalized_prompt as one undoable edit, while closing the dialog changes nothing. Generation never queues your sampler or save nodes - it prepares a prompt, not a render.

Inputs worth touching

  • agent - codex (default) or grok, set up under Settings → Arisu Nodes → Prompt Workbench → Agents: build the image, do the device-code login, pick a model and reasoning effort (only low, medium and high are offered).
  • skill - bundled:with-ref by default (Ref2VA / Hybrid: you have keyframes or references), bundled:no-ref for T2VA / I2VA / FL2VA / L2VA. Drop a folder with a SKILL.md at user/__arisu_nodes/skills/<skill-name>/ and it shows up as custom:<skill-name> - your house style, on demand.
  • requirements - the real brief, multiline. trigger_words - text you want present verbatim. motion_notes - what the previous clip did.
  • context_length - 5, 22, 39 or 56 frames, default 22; at 24fps that's about a second of previous clip. audio_context_length - 0–240, descriptive only; 0 follows the video length and no audio latent is decoded here.
  • Optional: video_settings and resources from the other Arisu H3 nodes, plus context_latent and vae - both required to enable Motion Context. prepare_job is the hidden job ticket; don't touch it.

Install

ComfyUI Manager → search ComfyUI-Arisu-Nodes → install → restart. Manual:

cd ComfyUI/custom_nodes
git clone https://github.com/swqa7697/ComfyUI-Arisu-Nodes.git

Needs ComfyUI ≥ 0.30.0 and Python ≥ 3.10; the pack is written against ComfyUI's V3 node API (io.ComfyNode / io.Schema). It declares two real runtime dependencies, Pillow and PyAV - check they landed in ComfyUI's interpreter, not your system Python. First startup writes user/__arisu_nodes/config.arisu.jsonc and a skills directory; the nodes live under Add Node → Arisu Nodes.

Agent generation additionally needs Docker with Linux container support that the ComfyUI process can reach, plus a policy permitting the CLI sandbox's nested user namespaces. No Docker, no agent: the left column disables itself and finalized_prompt stays editable, which beats a half-broken node.

Where people get burned

  • Expecting Queue to draft a prompt, or Generate to render a clip. Neither. Queue emits the text that's already there, and generation never queues downstream nodes.
  • Losing a draft. A new generation clears the previous output immediately, and activity plus unapplied output live outside node data - gone at shutdown. Apply it.
  • Motion Context greyed out. Both context_latent and vae are required, and arbitrary latent producers are rejected: a stock latent fails with expected H3 Motion Context Load Latent output.
  • Docker installed, generation still dead. It has to be the daemon the ComfyUI process can reach, with user namespaces allowed; when the sandbox can't be enforced, generation fails visibly rather than falling back to an unrestricted agent. Provider accounts are shared by everyone on that instance; credentials live in Docker volumes, not workflows.
  • The legal bit nobody puts in tutorials. The H3 weights are geofenced - the Community License excludes the US, EU, UK and South Korea. It's also 33B with ~42.5GB of reported full weights and no verified consumer floor.
CategoryArisu Nodes/MiniMax H3

Inputs (15)

NameTypeDefaultDescription
agentCOMBOcodex2 options: codex, grok
skillSTRINGbundled:with-ref—
context_lengthCOMBO224 options: 22, 5, 39, 56
audio_context_lengthINT240–240—
motion_notesSTRING—
reference_notesSTRING{}—
trigger_wordsSTRING—
requirementsSTRING—
finalized_promptSTRING—
video_settingsoptARISU_MINIMAX_H3_VIDEO_SETTINGS—
resourcesoptARISU_MINIMAX_H3_RESOURCES—
context_latentoptLATENT—
vaeoptVAE—
prepare_joboptSTRING—
motion_enabledoptBOOLEANtrue—

Outputs (1)

NameTypeDescription
promptSTRING—