Nodes/VELVET VICE — MiniMax H3/VELVET VICE MiniMax H3 — Prompt Director
ComfyUI Node

VELVET VICE MiniMax H3 — Prompt Director

Turn a source image and a one-line idea into a full H3 prompt

By Velvet-Vice·Created 10 days ago·Updated 7 days ago· 2
VELVET VICE MiniMax H3 — Prompt Director
  • image
  • prompt_package
  • final_prompt_preview
  • status
modeMANUAL
manual_prompt
short_idea
full_auto_settingsMODE: FULL_AUTO DURATION: AUTO — READ FROM H3 DIRECTOR DURATION CHOREOGRAPHY: AUTO — SAFE DEVELOPMENT BEATS SCENE COMPLEXITY: AUTO — DURATION BUDGET MODIFIER PARTICIPANT REGISTRY: AUTO — PRESERVE 1 TO 4 VISIBLE ADULTS ANATOMY CLASSIFICATION: WOMAN / FUTANARI / OTHER / UNCLEAR BEFORE ACTION CONTACT STATE MACHINE: ON GEOMETRY COMPOSER: LITERAL POSTURE + SUPPORT + PAIRWISE RELATIONS LEGACY FALLBACK: PRESERVE EXISTING PROVEN ACTION WHEN UNCERTAIN SCENE STYLE: passionate, explicit, physically natural and continuity-focused INTENSITY: PASSIONATE ACTION COMMITMENT: HIGH RHYTHM: STRONG CONTROLLED ESCALATION SOFT TISSUE PHYSICS: REALISTIC INERTIA: BODY-DRIVEN DAMPING: NATURAL RESPONSE SCALE: MOTION-MATCHED OSCILLATION LIMIT: SINGLE DAMPED SETTLE CONTACT DEFORMATION: ON CONTACT ANCHOR: SURFACE-LOCKED VOLUME PRESERVATION: ON GRAVITY RESPONSE: ON ENDING: CONTROLLED BY THE ENDING MODE SELECTOR CAMERA: PRESERVE VISIBLE NUDITY AUTO-TRIGGER: ON DIALOGUE: none unless clearly requested AUDIO: natural breathing, vocal reactions and environmental sound only; no music PERFORMANCE STYLE: LIFELIKE AND REACTIVE BODY LANGUAGE: WEIGHT-AWARE, ASYMMETRIC, CAUSAL FACIAL PERFORMANCE: CONTEXTUAL MICRO-REACTIONS MOTION TIMING: ORGANIC, NON-METRONOMIC AUDIO SYNC: ACTION-BOUND DIEGETIC SUPPORTING MOTION BUDGET: 1-2 RELEVANT CUES PER BEAT POSITION RECOGNITION: AUTO-DETECT FROM VISIBLE GEOMETRY POSITION MODE: PRESERVE CURRENT WHEN ACTIVE POSITION TRANSITION: MAXIMUM ONE LOW-COST TRANSITION ROLE ASSIGNMENT: GEOMETRY-BASED, ANATOMY-OWNER LOCKED MIXED ANATOMY RECOGNITION: STRICT LITERAL VISUAL EVIDENCE FUTANARI ANATOMY LOCK: PRESERVE BREASTS + ATTACHED PENIS PER OWNER PERSISTENT ANATOMY EXISTENCE: PRESENT_LOCKED SURVIVES OCCLUSION + RECEIVING ROLE OCCLUDED HAND COMPLETION: FORBIDDEN — PRESERVE VISIBLE SILHOUETTE ACTION CLASSIFICATION: SEPARATE EFFECTOR + TARGET FROM POSITION ACTION LOCK: PRESERVE MANUAL / ORAL / PENETRATIVE / OTHER CLASS POSITION VARIANT: AUTO-DETECT INDEPENDENTLY CONCURRENT ACTIONS: AUTO — ONE PRIMARY + ONE COMPATIBLE SECONDARY GENITAL REINTERPRETATION: FORBIDDEN The user has activated the separate graph-level confirmation that every depicted participant is 18 or older and consenting. The supported configuration may contain one to four adults already visible in the reference image with any visible adult anatomy or presentation. Treat the image as the literal first frame. Assign stable participants A through D as needed from visible screen position and identity cues; preserve their faces, hair, clothing, limbs, anatomy ownership, depth and support points. First resolve anatomy ownership, then detect the current position from independent geometry: each posture, facing direction, relative height, front/back order, pelvis relationship, support points and existing contact graph. Preserve an already active position instead of replacing it with a familiar alternative. If the position is neutral or uncertain, keep the visible postures and choose only the smallest physically reachable continuation. Assign initiator, receiver, active anatomy and target anatomy from existing contact, explicit request, visible compatibility, shortest continuous path, stable support and camera visibility—not from assumed gender. This includes solo scenes and any supported one-to-four-adult combination, including participants with similar or mixed visible anatomy. A feminine-presenting adult with breasts and a naturally attached visible penis is a valid futanari anatomy configuration: preserve the breasts and penis together, keep the penis attached to its original owner, and never replace it with a vulva-only body. For two or more futanari participants, resolve and lock each visible participant independently so every confirmed visible penis remains present. Treat anatomy existence separately from visibility: temporary body overlap, camera cropping, anal/vaginal receiving roles or an occluded pelvis never remove or relabel a `PRESENT_LOCKED` penis. A partially visible or occluded hand remains partial or occluded; never complete hidden fingers, add a sixth digit, duplicate a thumb or invent a new hand path. Never exchange, duplicate, erase, normalize or reinterpret anatomy. Independently classify the current or requested action from its active effector and exact target: hand-driven, mouth-driven, pelvis/anatomy-driven, body contact, visible object, preparation or unclear. A mounted/rider/cowgirl label describes position only and must never overwrite the action class. Preserve a confident manual, oral or other action exactly instead of converting it to a stereotypical alternative. Select one physically achievable primary adult action that best continues the image, with at most one compatible secondary action after resource validation. Use at most one low-cost position transition and only when duration and visible supports make every step achievable. Keep all existing image-fidelity, lifelike performance, physics, Ending Control and action-bound audio rules active.
adult_confirmedfalse
ollama_modelfredrezones55/Qwen3.5-Uncensored-HauhauCS-Aggressive:9b
ollama_urlhttp://127.0.0.1:11434
ollama_context_profile8-12 GB
ending_modeAUTO
prompt_profileMiniMax H3
duration_seconds5.0
ending_mode_override
audio_enabledtrue

Image-to-video prompting is where the blank-page problem gets brutal. You're not describing a still you can iterate on in seconds - you're describing motion over time, from a reference frame, and the prompt has to stay faithful to what's actually in that first image or the video drifts into something unrecognizable. The Velvet VICE MiniMax H3 Prompt Director is the node that does that planning for you: it looks at your source image, takes a rough idea, and produces the structured, duration-aware prompt package the rest of this H3 workflow consumes.

Be clear about the audience before you install it: this is an adult-content image-to-video pack, distributed through Civitai, and the Prompt Director's auto modes exist to script the kind of continuity-heavy explicit content that generic prompt enhancers mangle. If that's not your use case, the MANUAL path still works fine - but the vision modes and their default settings are built for NSFW work, and the node has a content-confirmation switch because of it.

How it works

Four modes, in escalating ambition:

  • MANUAL - you write the prompt. Crucially, this mode never contacts an LLM. Fully offline, works with just the pack installed, no external anything.
  • STANDARD VISION - the node sends your source image to a local vision LLM and gets back an analysis-driven prompt.
  • ADULT ASSISTED / ADULT FULL AUTO - the modes the pack is really about. The image and your short_idea go to a local uncensored model that plans a scene continuation: participant consistency across frames, anatomy ownership, position/geometry, action classification, escalation, and audio guidance, driven by a large editable settings block (full_auto_settings) that encodes the author's choreography rules.

The mechanism sits squarely in the "local LLM as a tool node in the graph" pattern. The director talks to an Ollama server on your own machine - default http://127.0.0.1:11434 - running an uncensored Qwen-family model. It encodes the image, sends the frame plus your idea plus the settings/profile, and parses the structured reply back into the pack's custom VELVET_VICE_PROMPT_PACKAGE. That's why the big modes need no API key: the filter that blocks explicit content lives in hosted APIs, and running the model locally is how you get around it. It's also why duration-aware planning works - the model is told the target duration and budgets the scene's beats to fit it.

Inputs that matter:

  • mode, manual_prompt, short_idea - described above.
  • adult_confirmed - the graph-level "every depicted person is 18+, consenting" gate. The auto modes expect it; keep it honest.
  • ollama_url / ollama_model / ollama_context_profile - point at your Ollama instance. The default model is a community uncensored Qwen variant ("9b"-class); swap in whatever you've pulled. The context profile (8–12GB / 16GB / 24–32GB+) tells the pack how to budget memory against your diffusion models.
  • ending_mode - AUTO / NO CLIMAX / CLIMAX / LOOP; H3's ending is owned by the separate Director, and this node defers to it.
  • Optional: image (the source frame), duration_seconds (up to ~149s), audio_enabled, plus an ending_mode_override and fixed prompt_profile.

Outputs: prompt_package (the custom type that feeds the H3 director/sampler chain), final_prompt_preview (a plain STRING you can route to a preview node or just read), and status.

Install - two pieces

The node, from the one-pack install:

cd ComfyUI/custom_nodes
git clone https://github.com/Velvet-Vice/velvet-vice-minimax-h3

Restart ComfyUI (or Manager → "velvet-vice-minimax-h3"). No pip dependencies.

And the brain, if you want any mode beyond MANUAL: an Ollama server running, with a model pulled:

ollama pull <the model name you put in ollama_model>

Manual mode never needs this. The other three die without it.

Troubleshooting, honestly

The first failure is "I set STANDARD VISION and nothing happened": check Ollama is actually running and that the model name in ollama_model is pulled - the node reports Ollama HTTP errors plainly when the server isn't there. Second: watch the preview output before committing to a render; a vision model can confidently describe motion the source image can't support, and the settings block is editable for a reason. And the security note that applies to every node in this category: it's arbitrary Python that shells out to a local HTTP server, so install it from the pack's GitHub/Registry and don't accept a modified copy from a stranger. Read the generated prompt once before you render a long clip - it's the cheapest bug fix in the workflow.

CategoryVELVET VICE/MiniMax H3

Inputs (14)

NameTypeDefaultDescription
modeCOMBOMANUAL4 options: MANUAL, STANDARD VISION, ADULT ASSISTED, ADULT FULL AUTO
manual_promptSTRING
short_ideaSTRING
full_auto_settingsSTRINGMODE: FULL_AUTO DURATION: AUTO — READ FROM H3 DIRECTOR DURATION CHOREOGRAPHY: AUTO — SAFE DEVELOPMENT BEATS SCENE COMPLEXITY: AUTO — DURATION BUDGET MODIFIER PARTICIPANT REGISTRY: AUTO — PRESERVE 1 TO 4 VISIBLE ADULTS ANATOMY CLASSIFICATION: WOMAN / FUTANARI / OTHER / UNCLEAR BEFORE ACTION CONTACT STATE MACHINE: ON GEOMETRY COMPOSER: LITERAL POSTURE + SUPPORT + PAIRWISE RELATIONS LEGACY FALLBACK: PRESERVE EXISTING PROVEN ACTION WHEN UNCERTAIN SCENE STYLE: passionate, explicit, physically natural and continuity-focused INTENSITY: PASSIONATE ACTION COMMITMENT: HIGH RHYTHM: STRONG CONTROLLED ESCALATION SOFT TISSUE PHYSICS: REALISTIC INERTIA: BODY-DRIVEN DAMPING: NATURAL RESPONSE SCALE: MOTION-MATCHED OSCILLATION LIMIT: SINGLE DAMPED SETTLE CONTACT DEFORMATION: ON CONTACT ANCHOR: SURFACE-LOCKED VOLUME PRESERVATION: ON GRAVITY RESPONSE: ON ENDING: CONTROLLED BY THE ENDING MODE SELECTOR CAMERA: PRESERVE VISIBLE NUDITY AUTO-TRIGGER: ON DIALOGUE: none unless clearly requested AUDIO: natural breathing, vocal reactions and environmental sound only; no music PERFORMANCE STYLE: LIFELIKE AND REACTIVE BODY LANGUAGE: WEIGHT-AWARE, ASYMMETRIC, CAUSAL FACIAL PERFORMANCE: CONTEXTUAL MICRO-REACTIONS MOTION TIMING: ORGANIC, NON-METRONOMIC AUDIO SYNC: ACTION-BOUND DIEGETIC SUPPORTING MOTION BUDGET: 1-2 RELEVANT CUES PER BEAT POSITION RECOGNITION: AUTO-DETECT FROM VISIBLE GEOMETRY POSITION MODE: PRESERVE CURRENT WHEN ACTIVE POSITION TRANSITION: MAXIMUM ONE LOW-COST TRANSITION ROLE ASSIGNMENT: GEOMETRY-BASED, ANATOMY-OWNER LOCKED MIXED ANATOMY RECOGNITION: STRICT LITERAL VISUAL EVIDENCE FUTANARI ANATOMY LOCK: PRESERVE BREASTS + ATTACHED PENIS PER OWNER PERSISTENT ANATOMY EXISTENCE: PRESENT_LOCKED SURVIVES OCCLUSION + RECEIVING ROLE OCCLUDED HAND COMPLETION: FORBIDDEN — PRESERVE VISIBLE SILHOUETTE ACTION CLASSIFICATION: SEPARATE EFFECTOR + TARGET FROM POSITION ACTION LOCK: PRESERVE MANUAL / ORAL / PENETRATIVE / OTHER CLASS POSITION VARIANT: AUTO-DETECT INDEPENDENTLY CONCURRENT ACTIONS: AUTO — ONE PRIMARY + ONE COMPATIBLE SECONDARY GENITAL REINTERPRETATION: FORBIDDEN The user has activated the separate graph-level confirmation that every depicted participant is 18 or older and consenting. The supported configuration may contain one to four adults already visible in the reference image with any visible adult anatomy or presentation. Treat the image as the literal first frame. Assign stable participants A through D as needed from visible screen position and identity cues; preserve their faces, hair, clothing, limbs, anatomy ownership, depth and support points. First resolve anatomy ownership, then detect the current position from independent geometry: each posture, facing direction, relative height, front/back order, pelvis relationship, support points and existing contact graph. Preserve an already active position instead of replacing it with a familiar alternative. If the position is neutral or uncertain, keep the visible postures and choose only the smallest physically reachable continuation. Assign initiator, receiver, active anatomy and target anatomy from existing contact, explicit request, visible compatibility, shortest continuous path, stable support and camera visibility—not from assumed gender. This includes solo scenes and any supported one-to-four-adult combination, including participants with similar or mixed visible anatomy. A feminine-presenting adult with breasts and a naturally attached visible penis is a valid futanari anatomy configuration: preserve the breasts and penis together, keep the penis attached to its original owner, and never replace it with a vulva-only body. For two or more futanari participants, resolve and lock each visible participant independently so every confirmed visible penis remains present. Treat anatomy existence separately from visibility: temporary body overlap, camera cropping, anal/vaginal receiving roles or an occluded pelvis never remove or relabel a `PRESENT_LOCKED` penis. A partially visible or occluded hand remains partial or occluded; never complete hidden fingers, add a sixth digit, duplicate a thumb or invent a new hand path. Never exchange, duplicate, erase, normalize or reinterpret anatomy. Independently classify the current or requested action from its active effector and exact target: hand-driven, mouth-driven, pelvis/anatomy-driven, body contact, visible object, preparation or unclear. A mounted/rider/cowgirl label describes position only and must never overwrite the action class. Preserve a confident manual, oral or other action exactly instead of converting it to a stereotypical alternative. Select one physically achievable primary adult action that best continues the image, with at most one compatible secondary action after resource validation. Use at most one low-cost position transition and only when duration and visible supports make every step achievable. Keep all existing image-fidelity, lifelike performance, physics, Ending Control and action-bound audio rules active.
adult_confirmedBOOLEANfalse
ollama_modelSTRINGfredrezones55/Qwen3.5-Uncensored-HauhauCS-Aggressive:9b
ollama_urlSTRINGhttp://127.0.0.1:11434
ollama_context_profileCOMBO8-12 GB3 options: 8-12 GB, 16 GB, 24-32+ GB
ending_modeCOMBOAUTO4 options: AUTO, NO CLIMAX, CLIMAX, LOOP / CONTINUOUS ACTION
imageoptIMAGE
prompt_profileoptCOMBOMiniMax H31 options: MiniMax H3
duration_secondsoptFLOAT5.01–149.5
ending_mode_overrideoptSTRING
audio_enabledoptBOOLEANtrue

Outputs (3)

NameTypeDescription
prompt_packageVELVET_VICE_PROMPT_PACKAGE
final_prompt_previewSTRING
statusSTRING