Nodes/SDXL Auto Prompter/APNext H3 Music Video (Minimal)
ComfyUI Node

APNext H3 Music Video (Minimal)

The one-box music video: song + lyrics + a cinematic look + three sliders (performance, pace, wildness) - the model invents the concept and the performer and writes the whole video. Attach reference images to fix the performer's face. The full H3 Music Video Writer runs underneath with sensible defaults; use that node when you need cast, locks, briefs or the masked-audio path.

By dagthomas·Created 3 years ago·Updated about 18 hours ago· 289
APNext H3 Music Video (Minimal)
  • audio
  • llm
  • image_1
  • image_2
  • image_3
  • image_4
  • image_5
  • image_6
  • image_7
  • image_8
  • image_9
  • scenes
  • durations
  • lengths
  • audio_segments
  • scenes_text
  • session_id
  • info
  • image_1
  • image_2
  • image_3
  • image_4
  • image_5
  • image_6
  • image_7
  • image_8
  • image_9
  • clip_starts
lyrics
visual_styleLive-action, 35mm cinematic film aesthetic
performance80
pace30
wildness45
modelsonnet
seed-1
Categorycomfyui_dagthomas/H3

Inputs (18)

NameTypeDefaultDescription
audioAUDIOThe song. It is cut into pieces on the music and every piece becomes one clip.
lyricsSTRINGLyrics, one line per line. Timestamps make the sync exact: `[0:15] line` (or LRC `[00:15.20] line`); section tags like [Chorus] are kept; untimed lines are spread evenly. Empty = instrumental video. The imagery of every scene is staged from its lyric lines.
visual_styleCOMBOLive-action, 35mm cinematic film aestheticThe look of the whole video - the curated cinematic looks (35mm, Wes Anderson, neon noir, ...) each fix style, camera, lenses and colour. Auto lets the model pick one to fit the song.
performanceINT800–100How much the singer is on camera. 0-33 = Narrative (story visuals, nobody sings on camera), 34-66 = Mixed (performance and story alternate), 67-100 = Performance (the singer lip-syncs the lyrics on camera).
paceINT300–100How fast the video cuts. 0 = long, slow pieces (up to ~15 s per clip), 100 = quick cuts (pieces down to ~6 s). The song is still cut ON the music inside that range.
wildnessINT450–1000 = grounded performance video, 100 = fully surreal. Above 40 seeds surreal events.
modelCOMBOsonnetWho writes the prompt. sonnet / opus / haiku / fable / default are Claude Code aliases (`default` = whatever the CLI is configured for). `codex` is the OpenAI Codex CLI with its configured model (shown when installed; `codex:<model-id>` in an H3 LLM Backend picks a specific one). ollama: / lmstudio: / local: entries are whatever your local servers were serving when the page loaded; pick one to run fully offline. Anything not listed goes in model_override.
seedINT-1-1–18446744073709550000Seeds the surreal picks and controls caching. -1 re-runs every queue.
llmoptAPNEXT_LLMOptional. Connect an APNext H3 LLM Backend node to write with Ollama, LM Studio, another OpenAI-compatible server or an API model instead of Claude Code. Overrides the model dropdown while connected.
image_1optIMAGEReference image 1: <Picture 1> in the prompt. Connect the same image to image_1 on the MiniMax H3 Reference to Video node, or use this node's image_1 output.
image_2optIMAGEReference image 2: <Picture 2> in the prompt. Connect the same image to image_2 on the MiniMax H3 Reference to Video node, or use this node's image_2 output.
image_3optIMAGEReference image 3: <Picture 3> in the prompt. Connect the same image to image_3 on the MiniMax H3 Reference to Video node, or use this node's image_3 output.
image_4optIMAGEReference image 4: <Picture 4> in the prompt. Connect the same image to image_4 on the MiniMax H3 Reference to Video node, or use this node's image_4 output.
image_5optIMAGEReference image 5: <Picture 5> in the prompt. Connect the same image to image_5 on the MiniMax H3 Reference to Video node, or use this node's image_5 output.
image_6optIMAGEReference image 6: <Picture 6> in the prompt. Connect the same image to image_6 on the MiniMax H3 Reference to Video node, or use this node's image_6 output.
image_7optIMAGEReference image 7: <Picture 7> in the prompt. Connect the same image to image_7 on the MiniMax H3 Reference to Video node, or use this node's image_7 output.
image_8optIMAGEReference image 8: <Picture 8> in the prompt. Connect the same image to image_8 on the MiniMax H3 Reference to Video node, or use this node's image_8 output.
image_9optIMAGEReference image 9: <Picture 9> in the prompt. Connect the same image to image_9 on the MiniMax H3 Reference to Video node, or use this node's image_9 output.

Outputs (17)

NameTypeDescription
scenesSTRING
durationsFLOAT
lengthsINT
audio_segmentsAUDIO
scenes_textSTRING
session_idSTRING
infoSTRING
image_1IMAGE
image_2IMAGE
image_3IMAGE
image_4IMAGE
image_5IMAGE
image_6IMAGE
image_7IMAGE
image_8IMAGE
image_9IMAGE
clip_startsFLOAT