Nodes/comfyui-music2video/🎵 Music2Video
ComfyUI Node

🎵 Music2Video

Turn one audio track into a finished music video — prompts first, money only if you want it

By lazniak·Created 3 days ago·Updated about 12 hours ago· 1
🎵 Music2Video
  • audio
  • reference_images
  • pipe
  • audio_clips
  • images
  • subject_images
  • videos
  • final_video
project_namemusic2video
iteration1
instructionCinematic music video. Describe the mood, story or visual world you want.
llm_providerlmstudio
lm_model
openrouter_model
openai_model
anthropic_model
visual_style
aspect_ratio16:9
clip_seconds6.0
min_shot_seconds5.0
max_shot_seconds15.0
num_shots0
creativity0.70
dynamicity0.60
word_influence0.6
whisper_deviceauto
seed0
image_providerpipe-steps
fal_image_model
fal_image_edit_model
openrouter_image_model
video_providerpipe-steps
fal_video_model
openrouter_video_model
render_concurrency2
lm_urlhttp://127.0.0.1:1234
lm_api_key
lm_model_override
lm_auto_downloadtrue
lm_auto_loadtrue
lm_context_length32768
lm_unload_aftertrue
lm_temperature0.80
lm_max_tokens4096
lm_timeout300
lm_retries2
lm_reasoning_effortnone
shots_per_request4
guide_excerpt_chars0
openrouter_api_key
openai_api_key
anthropic_api_key
fal_api_key
video_prompt_sourcei2va
render_subject_sheetsfalse
live_previewtrue
prompt_expansionminimal
lipsync_audiotrue
style_anchortrue
save_rendered_imagestrue
save_rendered_videotrue
concat_videotrue
final_audiomusic
final_fitpad
final_fps0
final_crf20
render_timeout600
whisper_modelopenai/whisper-large-v3
whisper_dtypefloat16
whisper_languageauto
whisper_chunk_length_s30
whisper_batch_size1
whisper_word_timestampstrue
whisper_window_seconds30
whisper_keep_loadedfalse
whisper_skipfalse
free_comfy_vramtrue
free_lmstudio_vramtrue
analyze_musictrue
snap_cuts_to_beatstrue
audio_clip_padding0.00
max_subjects6
negative_prompt_baseblurry, low quality, watermark, signature, text artifacts, deformed hands, extra limbs, oversaturated colors, jpeg artifacts, plastic skin
include_dialoguetrue
h3_style_directive
save_jsontrue
save_transcripttrue
filename_prefixmusic2prompts
verbosefalse
save_cost_reporttrue

This is the whole Music2Video pack in one node. You drop an audio track in, out comes everything you'd need to shoot a music video for it: shot-by-shot image prompts, MiniMax H3 video prompts in both of its formats, per-shot timings, negatives, and sample-accurate audio slices for lipsync - all on a single pipe. And if you want, it'll do the rendering too: start frames, clips, and the final cut-together film with the music muxed underneath. The name is not a lie about what it does; the surprise is which parts cost money. The analysis is always local and free, prompt writing is local by default, and rendering stays off until you pick a paid provider.

How it works

The pipeline is worth understanding because it's unusually honest about who does what. Whisper large-v3 transcribes the lyrics locally with word-level timestamps. librosa (when installed, otherwise a built-in numpy/scipy fallback) measures the BPM, beat grid, sections and energy. Then a shot planner in plain Python - no LLM touches timing - divides the track into 5–15 second shots snapped to the beats. An LLM writes the creative content: treatment, art direction, a "bible" of recurring subjects, then each shot. Finally, deterministic Python renderers assemble the exact MiniMax H3 prompt skeletons, so the format is correct no matter which model did the writing.

The LLM is never asked to "listen" or count time - it gets measured numbers and fixed shot boundaries and only invents the pictures and words. That's the pack's whole design thesis, and it's a good one: a small local model can write creatively without hallucinating a beat grid.

The inputs that matter

There are a lot of widgets here, but the 30+ advanced ones collapse into the node's native Advanced section (this is a V3-schema node, so that's a ComfyUI feature, not a hack). The ones you'll actually touch:

  • audio - the track, from LoadAudio. Its length defines the shot plan.
  • instruction - your brief: story, mood, world, constraints. It steers every writing stage.
  • llm_provider - lmstudio (local, free, default), or openrouter/openai/anthropic (billed per token).
  • clip_seconds, min_shot_seconds, max_shot_seconds - shot length. H3 only accepts 5–15 s, so the defaults respect that.
  • creativity, dynamicity, word_influence - three 0–1 sliders that shape how literal, how kinetic, and how lyric-following the writing is.
  • image_provider / video_provider - both default to pipe-steps, which renders nothing and bills nothing. Switch to fal or openrouter and it becomes a pay-per-call render farm.

Outputs

Six sockets. pipe (the M2P_PIPE type) carries every prompt, name, timing, the transcript and the analysis JSON - feed it to Music2Video Pipe Expand to get individual outputs. The media stay on their own sockets: images (start frames per shot), subject_images (reference sheets per subject), videos (per-shot clips), final_video (the assembled film), and audio_clips (each shot's exact slice of the track, at original sample rate - straight into a lipsync node).

Install

ComfyUI Manager: search Music2Video, install, restart. Or by hand:

cd ComfyUI/custom_nodes
git clone https://github.com/lazniak/comfyui-music2video music2video

It needs ComfyUI ≥ 0.3.48 (the first release with the V3 node schema it's built on). Dependencies are just requests, numpy, scipy, transformers, huggingface_hub - all already in a normal ComfyUI install. First run downloads Whisper large-v3 (~3 GB) to ComfyUI/models/whisper/. For local prompt writing, start LM Studio (Developer → Start Server, default http://127.0.0.1:1234); model dropdowns are fetched live, so if a model doesn't appear, right-click → Refresh model lists.

Where people get burned

VRAM. Whisper large-v3 takes ~2.9 GB of weights and up to ~7 GB peak during word-level timestamp alignment. On an 11 GB card that's the whole GPU. The node handles this aggressively: it unloads ComfyUI's models and the LM Studio model before transcribing, runs Whisper in 30-second windows, and steps down automatically (smaller batch → segment timestamps → CPU) if it still runs out of memory. Set whisper_device to a second card if you have one.

Reasoning models return empty. Gemma 4 and Qwen3 will spend the whole token budget on hidden thinking and hand back nothing. That's why lm_reasoning_effort defaults to none. If your stages come back empty, this is the first thing to check.

The billing surprise. One image per shot, one clip per shot. A 3-minute track cut into 6-second shots is 30 shots - 30 billed images plus 30 billed clips. The live_preview gallery (on by default) is there so you spot a bad prompt at shot 1 instead of paying for all of them, and the node writes a cost report to disk. Refreshingly upfront about the money.

Dead-looking runs. If you wire a media socket on a run where nothing rendered, the node deliberately hands you an execution blocker rather than an empty list - the downstream node is skipped and the log explains why, instead of the whole graph dying with a bare IndexError: list index out of range. Annoying the first time you see it; far better than the alternative.

CategoryMusic2Video

Inputs (84)

NameTypeDefaultDescription
audioAUDIOThe track. Its length is what the shot plan is divided out of; librosa analyses a mono copy at the original sample rate, Whisper gets a mono copy resampled to 16 kHz, and the per-shot slices on 'audio_clips' keep the original sample rate and channels. The same track is muxed under the finished film when 'final_audio' is 'music'.
project_nameSTRINGmusic2videoName of the folder this run writes into: ComfyUI/output/music2prompts/<project_name>_v<iteration>. One name per song or per client keeps a session's frames, clips, transcript, analysis JSON and cost report together instead of piled in one directory. It is one folder name, not a path: separators and anything Windows refuses become '-', so a name cannot write outside the output folder. Letters with diacritics are kept. Leave it empty and the folder is called 'music2video'.
iterationINT10–999999The take number, appended to the folder as _v001, _v002, ... The widget underneath is set to 'increment', so every run lands in its own folder and nothing overwrites the take you liked. Set it to 'fixed' to keep re-running into the same folder - the filenames still carry the timestamp, so even then nothing is overwritten. Anything above 999 simply widens: _v1000.
instructionSTRINGCinematic music video. Describe the mood, story or visual world you want.Your brief: story, mood, world, constraints. It is pasted verbatim into four stages - the interpretation, the art direction, the subject bible and every shot-content batch - so it steers the whole film. The image-prompt stage never sees it; what reaches the frames is whatever those stages made of it.
llm_providerCOMBOlmstudioWho writes the prompts. 'lmstudio' runs locally and is free; the other three bill per token, and one run is at least four requests plus two per batch of 'shots_per_request' shots. Switching away from 'lmstudio' hides that provider's server widgets ('lm_url', 'lm_api_key', 'lm_context_length', the auto-load switches) along with the other providers' model and key widgets, but 'lm_max_tokens', 'lm_timeout' and 'lm_retries' still apply to the cloud providers, and 'lm_temperature' and 'lm_reasoning_effort' apply to all of them except anthropic, whose request carries neither.
lm_modelCOMBOModel served by LM Studio, used only when 'llm_provider' is 'lmstudio'. The list is read from the running server; when the server was unreachable the dropdown falls back to a single built-in id ('google/gemma-4-e4b') that you may not have installed - start LM Studio, then right-click the node and choose 'Refresh model lists', or type the key into 'lm_model_override', which wins over this dropdown. With 'lm_auto_download' and 'lm_auto_load' on, the node installs and loads it itself.
openrouter_modelCOMBOModel used when 'llm_provider' is 'openrouter'; billed per token. The list is OpenRouter's public catalogue filtered to text-output models, so it fills in without a key. If you wire 'reference_images', pick one that accepts image input - the description stage sends them as images, and a model that refuses only produces a warning and no description.
openai_modelCOMBOModel used when 'llm_provider' is 'openai'; billed per token. Until a key is found in the environment the dropdown shows only a built-in fallback list ('gpt-5.2', 'gpt-5.1', 'gpt-5-mini', 'gpt-4.1'); the live catalogue - filtered to gpt / o1 / o3 / o4 / chatgpt chat ids - appears once OPENAI_API_KEY is set in the environment and you right-click the node and choose 'Refresh model lists'. The 'openai_api_key' widget is used for the run itself, not for this probe. Structured stages are sent as strict json_schema responses; the final retry drops the schema and parses the reply loosely.
anthropic_modelCOMBOModel used when 'llm_provider' is 'anthropic'. Structured stages go through forced tool use; temperature is not sent because current Claude models reject it, so 'lm_temperature' does nothing here. Until ANTHROPIC_API_KEY is set in the environment the dropdown shows only a built-in fallback list; the live catalogue appears after you set that variable and choose 'Refresh model lists' from the node's right-click menu. The 'anthropic_api_key' widget is used for the run, not for this probe.
visual_styleSTRINGForces a look, e.g. 'grainy 16mm night photography'. It reaches the art-direction stage only, as a hard requirement; empty lets that stage choose. Every later stage sees the art direction it produced, not this text, so a short entry gets elaborated rather than pasted into the prompts.
aspect_ratioCOMBO16:9Framing for the image prompts and, when rendering, for the payload: fal endpoints get a pixel size derived from it (1024-based, snapped to multiples of 32) or the nearest named preset, OpenRouter gets the string as it stands. A value the endpoint does not declare is dropped rather than substituted, and on preset-only endpoints 21:9 collapses to landscape_16_9. On a clip that goes out with a start frame, the aspect is not sent at all - the frame decides it.
clip_secondsFLOAT6.01–60Target length of one shot before pacing: 'dynamicity' scales it from 1.3x (at 0) to 0.7x (at 1), the result is clamped into 'min_shot_seconds'..'max_shot_seconds', and the shot count is the track length divided by that. It is ignored completely when 'num_shots' is anything other than 0.
min_shot_secondsFLOAT5.00.5–60Shortest shot the planner will produce. It also caps the shot count at track length / this value, so it silently overrules a larger 'num_shots', and a track shorter than it comes back as one single shot. Shots under 2 s are refused by every audio field this node can fill, so their slice goes out without the vocal.
max_shot_secondsFLOAT15.01–60Longest shot allowed: the planner adds shots until none is longer, then splits any span that still exceeds it. Each shot's duration is what the video endpoint is asked for, clamped to the range or enum that endpoint declares - so a shot longer than the model allows comes back short, and 'concat_video' holds its last frame to fill the gap. Past 15 s a shot also loses its audio slice on endpoints that take reference audio as a list field, so lipsync stops for it.
num_shotsINT00–4000 derives the count from the track length, 'clip_seconds' and 'dynamicity'. Any other value asks for exactly that many and makes 'clip_seconds' irrelevant, but it is still clamped: never more than track length / 'min_shot_seconds', and shots are added back if that would push any of them past 'max_shot_seconds'. Every shot is one LLM shot-content entry and, when rendering, one paid image and one paid clip.
creativityFLOAT0.700–10 = grounded and literal, 1 = bold and surreal. It goes to the art-direction stage alone, printed into its brief as a number; it is not a sampling parameter - 'lm_temperature' is the one that changes decoding - and it reaches the shots only through the art direction it produced.
dynamicityFLOAT0.600–1Two effects: it scales 'clip_seconds' by 1.3 at 0 down to 0.7 at 1 while the shot count is being derived, and it is printed into every shot-content request as the pacing level (0 = calm and static, 1 = restless and kinetic). The scaled value is clamped into 'min_shot_seconds'..'max_shot_seconds' first, so at the defaults (6 s target, 5 s floor) anything above about 0.78 no longer shortens the shots - only the pacing wording changes. With 'num_shots' set, only the wording effect remains.
word_influenceFLOAT0.6-1–1+1 = visualise the lyrics literally, -1 = ignore the words and use the vibe. The interpretation stage only sees three bands - above +0.33, below -0.33, and everything between reads identically - while the exact number is printed into every shot-content request. With 'whisper_skip' on or an instrumental track there are no words to weight.
whisper_deviceCOMBOautoWhere Whisper runs. 'auto' takes cuda:0 when CUDA is present and CPU otherwise; a cuda index that does not exist falls back to cuda:0 with a warning, and CPU forces float32 whatever 'whisper_dtype' says. Pointing it at a second card is the alternative to the unloads run before transcription - 'free_comfy_vram', and 'free_lmstudio_vram' when the LLM is LM Studio.
seedINT00–11529215046068470000 means no seed is sent at all - not seed zero - so the LLM and every render run unseeded. Any other value is applied per item, not per run: shot 1 gets the seed itself, shot 2 seed+1 and so on; the first subject sheet gets seed+1000; each clip uses the same offset as its shot. Everything is folded into the 32-bit range providers accept (abs(seed) modulo 2^31), and fal folds it again into whatever range the endpoint declares, so distant seeds can land on the same value. The LLM seed reaches LM Studio, OpenAI and OpenRouter; the Anthropic request carries no seed field, so on that provider the prompt writing is unseeded whatever you set here.
image_providerCOMBOpipe-stepsWhere the start frames are rendered. 'pipe-steps' skips the images: nothing is billed for them and the shots go out as prompts only. 'fal' and 'openrouter' are paid per-call APIs - one image per shot, plus one per subject when 'render_subject_sheets' is on. Setting it also unhides the render widgets, and a fal video model whose schema requires an image aborts the run at once while this is 'pipe-steps', before anything is billed. Neither this switch nor 'video_provider' affects the LLM: a cloud 'llm_provider' still bills per token.
fal_image_modelCOMBOThe plain text-to-image fal endpoint. It always draws the subject sheets, and it draws the shots as well whenever there is nothing to reference (no subject sheets, no 'reference_images'); once references exist the shots move to 'fal_image_edit_model'. The list is fal's text-to-image and image-to-image index combined, so edit endpoints appear in it too - one picked here has nothing to edit unless you wired 'reference_images'.
fal_image_edit_modelCOMBOfal image model used for the shots once there is an identity to hold on to - the subject sheets, or your 'reference_images'; with neither present it is never consulted and 'fal_image_model' draws everything. The style anchor is added to what this model receives, but it never selects it on its own. This has to be an endpoint that accepts a reference image: the node reads the endpoint's schema and warns when it declares no field for one, and a plain text-to-image model silently throws the identity away so every shot comes back a different person. Edit endpoints usually carry 'edit', 'kontext' or 'image-to-image' in the id, but the schema, not the name, is what decides.
openrouter_image_modelCOMBOImage model used when 'image_provider' is 'openrouter'. The node always sends the subject sheets and the style anchor as 'input_references' on this provider - there is no OpenRouter counterpart to 'fal_image_edit_model' and no capability check, so a model that ignores references simply loses the identity. The request carries prompt, aspect ratio, seed and references: negatives ('negative_prompt_base' and the per-shot extras) are not part of that API and are dropped.
video_providerCOMBOpipe-stepsWhere the clips are rendered. 'pipe-steps' skips the clips: nothing is billed for them. 'fal' and 'openrouter' are paid per-call APIs. Only 'fal' is checked against the endpoint's published schema before submitting - the OpenRouter payload is fixed (prompt, duration, aspect ratio, seed, and either a first frame or references), and it has no input for a driving audio track at all, so 'lipsync_audio' cannot work there. Neither this switch nor 'image_provider' affects the LLM: a cloud 'llm_provider' still bills per token.
fal_video_modelCOMBOVideo model used when 'video_provider' is 'fal'. Its published schema decides most of the payload: the shot duration is clamped to the range or enum it declares, and 'lipsync_audio' only does something if it declares a real audio input. The first-frame/reference choice comes from 'video_prompt_source' plus the id - with 'ref2va' selected, an id containing 'reference-to-video' is sent the subject references instead of a first frame. If the schema marks an image field required while no images are being rendered, the run stops before anything is billed.
openrouter_video_modelCOMBOVideo model used when 'video_provider' is 'openrouter'. The payload is fixed and unverified - prompt, duration rounded to whole seconds, aspect ratio, seed, and either a first frame or references - so a model that expects anything else simply fails. 'lipsync_audio' and 'prompt_expansion' have no field to go into on this API and are ignored.
render_concurrencyINT21–16How many renders are in flight at once within one pass - the subject sheets, then the start frames, then the clips - never across passes. 1 runs them in a plain loop; when 'style_anchor' is on and there is something to reference, shot 1 is rendered alone first whatever this is set to. The limit that bites is your provider's own concurrency allowance rather than the 16 here. Failures are per item: a failed image keeps its slot as a black frame, a failed clip is dropped from 'videos' and from the final cut.
lm_urlSTRINGhttp://127.0.0.1:1234Address of the LM Studio server: the model list, the load/unload calls and the completions all go here. Only used when 'llm_provider' is lmstudio - the three cloud providers have fixed endpoints and this widget is hidden for them. Default http://127.0.0.1:1234. Nothing here is contacted until the model-preparation step, which runs after transcription, so an unreachable server fails the run minutes in with 'LM Studio unreachable' rather than immediately (start it under Developer -> Start Server).
lm_api_keySTRINGBearer token for LM Studio, sent only when this field is non-empty. LM Studio needs none by default - fill it in only if you put the server behind a proxy that demands one. Unlike the four cloud key fields, this one has no environment-variable fallback: empty means no Authorization header at all.
lm_model_overrideSTRINGModel key sent instead of the 'lm_model' dropdown, e.g. google/gemma-4-e4b. Use it when LM Studio was offline while ComfyUI built the dropdown, so the model you want was never in the list. Ignored for every provider other than lmstudio - openrouter/openai/anthropic always use their own dropdown.
lm_auto_downloadBOOLEANtrueWhen the chosen key is not installed in LM Studio yet, ask the server to download it. If LM Studio returns a job id the node waits for it and reports progress at most every 5 s, capped at max(600 s, 4x 'lm_timeout'), after which the run continues anyway; if it returns no job id the node does not wait at all. Off = the node only warns that the model is missing and skips the whole load/reload step. lmstudio only.
lm_auto_loadBOOLEANtrueLoad the model before the stages start, and reload it when the instance LM Studio currently holds was loaded with a smaller context than 'lm_context_length'. Off = whatever is loaded is used as it stands, including a context too small for this node's prompts. If the load call fails, the node retries without the context setting and then falls back to LM Studio's just-in-time loading. lmstudio only.
lm_context_lengthINT327684096–262144Context the model is loaded with, in tokens. If it is already loaded with less, the node unloads and reloads it at this size - but only while 'lm_auto_load' is on. This is the input side: 'guide_excerpt_chars' adds its characters to every shot-content request and 'shots_per_request' decides how many shots share one request, so raise this when you raise either. 'lm_max_tokens' caps the reply, not the prompt.
lm_unload_afterBOOLEANtrueUnload the model from LM Studio once the prompts are written, before any image or video rendering starts, and wait until LM Studio reports it gone so the VRAM is really back. On by default: a local LLM left resident is the usual reason the sampler or the video model that runs next cannot find room. Turn it off to keep the model warm when this node is the only thing in the graph. lmstudio only: the cloud clients' unload is a no-op.
lm_temperatureFLOAT0.800–2Sampling temperature applied to every LLM stage, from the track interpretation to the image prompts. Sent to LM Studio, OpenAI and OpenRouter. Never sent to Anthropic - current Claude models reject the field - so on 'llm_provider' = anthropic this widget does nothing.
lm_max_tokensINT4096256–32768Ceiling on one stage's reply, per request - not per run, and one request has to hold every shot in its batch. Run out and the request fails: either with a JSON parse error on the truncated reply, or, when the model wrote nothing at all, with 'the model hit the token limit before answering'. A failed shot-content or image-prompt batch does not stop the run - those shots come back with empty content, and a missing image prompt is rebuilt from the shot text and the art direction - so raise this or lower 'shots_per_request'. On the paid providers it is also the cap on billed output tokens per request.
lm_timeoutINT30030–3600Seconds one LLM request may take before it is abandoned; the widget's own minimum is 30. It also sizes the lifecycle calls - a model load gets max(120 s, this) and a download waits up to max(600 s, 4x this). A timeout counts as a failed attempt and consumes one of 'lm_retries'.
lm_retriesINT20–5Extra attempts per stage after a failure, with a growing wait between them (1.5 s times the attempt number, so 1.5 s then 3 s at the default of 2). On the final attempt the JSON schema is dropped and the model is simply told to reply with raw JSON - the fallback that rescues models which cannot do structured output, and which you lose entirely at 0. Every retry is a fresh billed request on the paid providers.
lm_reasoning_effortCOMBOnoneThinking budget sent with each request. 'none' is the default because a reasoning model (Gemma 4, Qwen3) otherwise spends the whole 'lm_max_tokens' budget inside reasoning_content and returns an empty message, which fails the stage. LM Studio is sent every value except 'default' and retries without the field if the model rejects it; OpenAI and OpenRouter only receive low/medium/high; Anthropic never receives it.
shots_per_requestINT41–16How many shots go into one LLM request, in both the shot-content stage and the image-prompt stage: 20 shots at 4 means 5 requests each. Lower is safer for small models because less has to fit in 'lm_max_tokens', at the cost of more round trips. In the shot-content stage only, each batch after the first is also shown the previous batch's last shot (first 1200 characters) so the look carries across the seams.
guide_excerpt_charsINT00–40000Characters of the official MiniMax H3 guides pasted into the shot-content system prompt. The budget is spent in order: base-en.txt takes up to half of it, then ref-en.txt takes up to half of what is left, so base gets about twice as much as ref and roughly three quarters of the number you type is actually injected. It only does something when ComfyUI-MiniMaxH3-Easy sits beside this pack in custom_nodes; if it does not, nothing is injected and nothing is logged. The text is repeated in every shot-content request, so raise 'lm_context_length' with it. 0 (default) sends the compact built-in rules only.
openrouter_api_keySTRINGKey for OpenRouter - one field covers all three uses: the LLM stages ('llm_provider' = openrouter) and the image and video rendering ('image_provider' / 'video_provider' = openrouter). Empty falls back to OPENROUTER_API_KEY, then OPEN_ROUTER_API_KEY, then OPENROUTER_KEY; a value typed here wins over all of them. The 'openrouter_model' list is public and fills in without any key.
openai_api_keySTRINGKey for the OpenAI LLM stages; OpenAI is never used for image or video rendering here. Empty falls back to OPENAI_API_KEY, then OPEN_AI_API_KEY; a value typed here wins over both. The 'openai_model' dropdown is probed with the environment variables only, so a key typed here authorises the run but leaves that list at its built-in fallbacks.
anthropic_api_keySTRINGKey for the Anthropic LLM stages; Anthropic is never used for image or video rendering here. Empty falls back to ANTHROPIC_API_KEY, then ANTHROPIC_AUTH_TOKEN; a value typed here wins over both. The 'anthropic_model' dropdown is probed with the environment variables only, so a key typed here authorises the run but leaves that list at its built-in fallbacks.
fal_api_keySTRINGKey for fal.ai, used for rendering images and clips only - fal is not one of the LLM providers. Empty falls back to FAL_KEY, then FAL_API_KEY, then FAL_ADMIN_API_KEY; a value typed here wins over all three. With no key anywhere the run stops the moment rendering starts ('no fal.ai key'); the fal model dropdowns come from fal's public index and need none.
video_prompt_sourceCOMBOi2vaWhich prompt set is actually sent to the video model. Both are always written to the 'video_prompts_i2va' and 'video_prompts_ref2va' outputs - this only picks the one that gets rendered. 'i2va' sends the rendered start frame as the first frame. 'ref2va' sends references instead of a start frame - the wired reference_images and the rendered subject sheets, in that order, capped at the first 9. It needs two things: at least one reference (turn on 'render_subject_sheets' with an 'image_provider', or wire reference_images), and a video endpoint that declares a reference field, such as minimax/h3/reference-to-video. Without references the node warns and the clip goes out as text only; on an endpoint whose schema has only a first-frame field, just the first reference is sent as that frame; and a fal endpoint that requires an image aborts the run before anything is billed.
render_subject_sheetsBOOLEANfalseRender one reference image per subject (up to 'max_subjects', default 6) before the shots, billed per image like any other. On fal they are drawn by the plain 'fal_image_model', because an edit model cannot draw a subject that does not exist yet - and once they exist the shot frames switch over to 'fal_image_edit_model'. On OpenRouter there is no switch: one image model handles both. Not strictly required for 'video_prompt_source' = ref2va - wired reference_images are sent too - but only rendered sheets get the <Picture N> number the prompt cites, so without them the references go out unlabelled.
live_previewBOOLEANtruePush each finished image and clip into the node's gallery the moment it lands, instead of after the whole batch - which is how you catch a bad prompt at shot 1 rather than paying for twelve. It costs nothing: files go to ComfyUI/temp/music2prompts, and when 'save_rendered_video' is off the clips already written there are reused instead of written twice. The gallery does not survive a page reload; the results themselves come back on the IMAGE and VIDEO outputs.
prompt_expansionCOMBOminimalHow much the video endpoint may rewrite the prompt before generating. 'minimal' and 'rich' set whichever of the two fields the endpoint declares: 'prompt_expansion_mode' to 'fast' or 'quality' (the MiniMax H3 endpoints), 'enable_prompt_expansion' to false or true (fal-ai/wan/v2.7/image-to-video). 'model default' sends neither. On an endpoint that declares neither - and every OpenRouter video model, since that payload has no such field at all - this setting does nothing. It matters because every MiniMax H3 endpoint defaults to prompt_expansion_mode 'balanced', which decides per request, so each shot's look is re-invented independently of the art direction.
lipsync_audioBOOLEANtrueSend each shot's own slice of the track (MP3, inline in the request) so the performance follows the vocal. It only reaches endpoints that declare an audio input: on fal that is audio_url (fal-ai/wan/v2.7/image-to-video) or reference_audio_urls (minimax/h3/reference-to-video); a boolean named 'audio' means 'generate a soundtrack' and is skipped, and OpenRouter's video API has no audio input at all - the node warns instead of sending. The accepted window is 2-15 s for the list-shaped reference field and 2-30 s for a single audio_url; a shot outside it is warned about and rendered without audio, so watch 'max_shot_seconds' and 'audio_clip_padding', which widens every clip.
style_anchorBOOLEANtrueRender shot 1 on its own first, then hand it to every later shot as an extra reference - this is what holds the grade, the grain and the wardrobe together. The cost is serialisation: that one image renders alone while 'render_concurrency' is ignored, and only the remaining shots run concurrently. It does nothing unless there are at least 2 shots and at least one reference already in play (wired reference_images or 'render_subject_sheets'), and on fal it needs a frame model that declares a reference field, i.e. 'fal_image_edit_model'. If shot 1 fails, the rest go out without an anchor.
save_rendered_imagesBOOLEANtrueKeep the rendered start frames and subject sheets in ComfyUI/output/music2prompts, named <prefix>_<stamp>_frame001 and <prefix>_<stamp>_subject001, in whatever format the endpoint returned (png/jpg/gif/webp is detected from the bytes; anything unrecognised is written with a .png name). Off only skips that write - the renders were paid for either way, still leave the node on the 'images' and 'subject_images' outputs, and with 'live_preview' on a copy of each is still written to ComfyUI/temp/music2prompts for the gallery. Hidden while both 'image_provider' and 'video_provider' are 'pipe-steps'.
save_rendered_videoBOOLEANtrueThe clips are always written to disk - the 'videos' output and the final film both need real files - so this only picks the folder: on, ComfyUI/output/music2prompts as <prefix>_<stamp>_shot001.mp4; off, ComfyUI/temp/music2prompts, which ComfyUI clears out, and there the file is normally the one 'live_preview' already wrote, named <prefix>_<stamp>_video001.mp4. The concatenated film always lands in the output folder regardless.
concat_videoBOOLEANtrueRe-encode the clips into one H.264/mp4 on the 'final_video' output, written to ComfyUI/output/music2prompts as <prefix>_<time>_final.mp4 (PyAV, no ffmpeg binary, and never a stream copy). Every clip is re-timed onto one grid to the exact length of its shot: one that came back long is cut, one that came back short holds its last frame so the film stays in sync with the music. Shots whose render failed are left out entirely, so the film ends up shorter than the track by their length.
final_audioCOMBOmusic'music' muxes the track you fed in as AAC, cut to the film's length or padded with silence if the film outlasts it. 'clips' keeps the audio the video model returned, padding any silent clip so the cuts stay aligned, and drops to silence if no clip carries an audio track at all. 'none' leaves the film silent.
final_fitCOMBOpadWhat happens to a clip whose aspect differs from the film's: 'pad' letterboxes it on black, 'crop' scales up and cuts the edges off centre, 'stretch' distorts it to fill the frame. A clip whose aspect ratio is within 0.005 of the film's is only scaled, so this setting does nothing when every clip has the same shape - which is the normal case, since all shots are rendered at one 'aspect_ratio'.
final_fpsFLOAT00–120Frame rate of the finished film; every clip is resampled onto it by duplicating or dropping frames. 0 takes the highest rate any clip reports - so slower clips get frames duplicated rather than faster ones losing them - ignoring rates above 120 fps as mis-reported and falling back to 24 if no clip reports one. A value near 23.976 / 29.97 / 59.94 is snapped to the exact 1001-based rational.
final_crfINT200–51libx264 -crf for the final film, over the widget's 0-51 range: lower is better quality and a bigger file, 20 is the default, and the encode runs at preset 'medium'. It applies to the concatenated film only - the individual clips are stored exactly as the provider returned them.
render_timeoutINT60060–3600Seconds one image or clip may take before the node gives up on it; that shot then comes back empty - a black placeholder frame keeps its slot on 'images', a failed clip is simply missing - while the rest of the batch continues. It bounds the fal queue polling (checked every 2 s) and the OpenRouter image request and video polling (every 3 s); the submit call and the download of the finished file have their own fixed timeouts. It has nothing to do with 'lm_timeout' - the LLM stages are timed separately.
whisper_modelCOMBOopenai/whisper-large-v3Which local Whisper does the transcription - it runs on this machine, nothing is uploaded and nothing is billed. On first use the weights are pulled from HuggingFace into ComfyUI/models/whisper/<repo--id>, several GB, once per model. Measured for large-v3 on an 11 GB card: about 2.9 GB for the weights, ~7 GB peak with word timestamps and ~4 GB with segment timestamps.
whisper_dtypeCOMBOfloat16Precision of the Whisper weights. float16 halves the VRAM of float32 and is chosen deliberately over bfloat16 so Turing cards stay supported. It applies on GPU only: if 'whisper_device' resolves to cpu, or no CUDA is present, the run is forced to float32 and this widget does nothing.
whisper_languageSTRINGauto'auto' lets Whisper detect the language; anything else is passed to the decoder as the forced language. Whatever you type here is also used verbatim as the [Language] tag inside the H3 <d>...</d> dialogue blocks, so 'Polish' reads better in a prompt than 'pl'. On 'auto' that tag comes from a character-set guess that can only tell apart English, Polish, German, Spanish, Chinese and Japanese - anything else is tagged English.
whisper_chunk_length_sINT305–30Length of the audio chunks the transformers ASR pipeline decodes when the input is longer than one chunk; 30 is both the default and the widget's maximum. It does not bound the memory word timestamps need - that grows with the whole input, which is what 'whisper_window_seconds' is for. Lowering it does not fix an out-of-memory error.
whisper_batch_sizeINT11–32How many chunks are decoded at once. Word-level timestamps peak around 7 GB of VRAM on an 11 GB card even at 1, so raise this only on a card with memory to spare. If an attempt fails - out of memory or for any other reason - the run steps down a ladder by itself: batch 1, then segment timestamps instead of word ones, then no timestamps at all (which leaves every shot without lyrics), then CPU at float32.
whisper_word_timestampsBOOLEANtrueAsk for per-word timings - a DTW pass over the cross-attentions - so each lyric is assigned to the shot its midpoint falls in. This is the expensive part of transcription: ~7 GB peak against ~4 GB for segment-level timings on an 11 GB card. Off, words are still timed but only per segment, so a line can be attributed to the neighbouring shot.
whisper_window_secondsFLOAT300–600Transcribe the track in windows of this many seconds, shifting each window's timings back into track time. Word-timestamp memory grows with the length of the whole input, not with 'whisper_chunk_length_s' - measured on an 11 GB card at ~7 GB for 60 s and an out-of-memory failure at 90 s - hence the 30 s default. 0, or any value below 5, disables windowing and sends the whole track in one pass; windowing is also skipped when the track is shorter than the window plus 5 s.
whisper_keep_loadedBOOLEANfalseKeep the Whisper pipeline in memory after the run - keyed by model, device and dtype - so the next queue does not reload several GB from disk. Off by default, because large-v3 holds about 3 GB and the samplers downstream usually want the whole card; turn it on when you run this node repeatedly and have the headroom. Off unloads it and empties the CUDA cache as soon as the transcript is done, freeing that VRAM for the rest of the workflow at the price of a full reload next time. Only one pipeline is cached, so changing 'whisper_model', 'whisper_device' or 'whisper_dtype' replaces it anyway.
whisper_skipBOOLEANfalseSkip transcription entirely, for an instrumental track. The 'transcript' output is then empty, no lyrics reach any shot, the model is told to leave every dialogue field empty whatever 'include_dialogue' says - so the H3 <d> blocks normally disappear - and the H3 language tag falls back to English. 'free_comfy_vram' and 'free_lmstudio_vram' also never run, because nothing needs the VRAM.
free_comfy_vramBOOLEANtrueCall ComfyUI's unload_all_models() and empty its cache just before Whisper loads, so a checkpoint another node left resident cannot push the transcription into an out-of-memory error. Those models reload the next time they are used. Does nothing when 'whisper_skip' is on.
free_lmstudio_vramBOOLEANtrueUnload the LM Studio model before Whisper starts; it is loaded again three steps later, at the 'preparing model' step just before the writing passes, which needs 'lm_auto_load' on - with that off the node never loads it back itself. Keep this on with a single GPU: an 11 GB card cannot hold the LLM and Whisper large-v3 (~2.9 GB of weights, ~7 GB peak) at once. Shown only for llm_provider 'lmstudio', and skipped when 'whisper_skip' is on.
analyze_musicBOOLEANtrueMeasure tempo, the beat grid, section boundaries and an energy curve, and hand them to the LLM as facts instead of asking it to imagine the music. librosa does this when it is installed; otherwise a numpy/scipy fallback runs and the analysis JSON records which one under 'backend'. Off, the track becomes one section called 'Part 1' at 0 BPM with no beats - which also leaves 'snap_cuts_to_beats' nothing to snap to.
snap_cuts_to_beatsBOOLEANtrueMove each shot boundary onto the nearest section edge, or onto the nearest beat when no section edge is close enough, so cuts land on the music rather than on an even division of the track. A boundary moves at most 42% of an even division onto a section edge and 35% onto a beat, and only when the move still leaves every shot at or above 'min_shot_seconds'. Does nothing with 'analyze_music' off: there is then no beat grid, and the single section spans the whole track, so no boundary has anything to move to.
audio_clip_paddingFLOAT0.000–2Widen every shot's audio clip by this many seconds on both sides, clamped to the track. It always affects the 'audio_clips' output; it affects the slice sent to a video model as driving audio only when 'lipsync_audio' is on, the video provider is fal and the chosen fal model declares an audio field (OpenRouter's video API has none). 0 keeps the cut sample-accurate against the shot boundaries, which is what lipsync needs - any padding puts the vocal out of step with the frames. Padding also lengthens the clip, and a clip outside the accepted window (2-15 s for a reference-audio list such as H3, 2-30 s for a driving-audio field) is sent as no audio at all.
max_subjectsINT60–16Upper bound on the recurring characters, locations and props locked for consistency: the model is told to write at most this many and the list is then truncated to it. With 'render_subject_sheets' on, each subject costs one paid reference image. Those sheets reach the video model only with 'video_prompt_source' 'ref2va', and then only the first 9 references (wired 'reference_images' plus sheets) are sent, with the prompt's <Picture N> labels numbered against that same list; on the default 'i2va' the sheets are used as references for the start frames instead. 0 leaves no subjects, so no sheets are rendered and a ref2va prompt has nothing to point at.
negative_prompt_baseSTRINGblurry, low quality, watermark, signature, text artifacts, deformed hands, extra limbs, oversaturated colors, jpeg artifacts, plastic skinMerged with the art direction's own negative terms and each shot's - split on commas and semicolons, dots trimmed off the ends, de-duplicated case-insensitively - into the 'negative_prompts' output. It only reaches fal image endpoints that declare a 'negative_prompt' field, or one whose schema could not be read, where it is sent blind; either way it is the first field dropped when the endpoint rejects the payload. OpenRouter's image API has no negative field and no video path sends one at all, so on those it is written to the output and otherwise ignored.
include_dialogueBOOLEANtruePut each shot's transcribed lyrics into the H3 <d>[Language] ...</d> block, rendered as '<subject> (S1) sings: ...' - that block is what makes a lipsync-capable model perform the line. Off, the per-shot lyrics are withheld from the shot-writing stage and the model is told to leave the dialogue field empty, so no <d> block appears in either the i2va or the ref2va prompt; the full transcript still goes to the first stage, which reads the track before any shot is written. When Whisper returned no words (including under 'whisper_skip') the model gets that same instruction, so in practice no <d> block is written - but nothing downstream strips one if the model writes a line anyway.
h3_style_directiveSTRINGExtra clause appended to the style line of every H3 prompt, in both the i2va and the ref2va form - for example '35 mm film grain, no lens flare'. It is joined to the LLM's own visual_style with a comma, so write attributes, not sentences; a leading field label such as 'style:' is stripped and a trailing full stop removed. Empty leaves the style line exactly as the art direction wrote it.
save_jsonBOOLEANtrueWrite the whole run to ComfyUI/output/music2prompts/<prefix>_<stamp>_analysis.json: the music analysis, the transcript, the interpretation and art direction, every subject, per shot its times, section, lyrics and the raw content the model wrote, plus which providers and models were used for rendering, how many images and sheets came back, and the file path of every clip. The same text is on the 'analysis_json' output whether this is on or off.
save_transcriptBOOLEANtrueWrite <prefix>_<stamp>_transcript.txt next to the JSON: the detected language, the shot count, the full transcript, then one block per shot with its time range, section, the image prompt it was rendered from, and the first line of its i2va video prompt (capped at 200 characters), and the words sung inside that shot. A shot with no words reads '(instrumental)'. The 'transcript' output carries only the bare Whisper text, with no timings and no shot structure.
filename_prefixSTRINGmusic2promptsLeading part of every file this run writes. One timestamp is taken when the run starts and shared by the frames, subject sheets, clips, transcript and JSON (<prefix>_<stamp>_frame001, _subject001, _shot001.mp4, _transcript.txt, _analysis.json), so a run's files sort together; the concatenated film takes a fresh timestamp at the moment it is assembled, as <prefix>_<time>_final.mp4. Empty falls back to 'music2prompts' for the transcript, the JSON and the final film only - the frames, sheets and clips get no fallback and are written with a leading underscore instead.
verboseBOOLEANfalseLog the number of characters each LLM reply came back with - which is how a truncated reply is told apart from a refused one. On LM Studio the line names the stage; the OpenAI-compatible providers name the provider instead, so those lines are only told apart by their order, and Anthropic logs nothing for a stage that comes back through its structured tool call. The stage-by-stage progress lines, the render counts and every warning are printed regardless of this switch. It is also handed to the fal and OpenRouter media clients, where it currently produces no extra output.
save_cost_reportBOOLEANtrueWrite <prefix>_<stamp>_cost.json and _cost.txt next to the JSON: every billed call of the run with its model, what it billed and how the price was arrived at, plus the per-model subtotals and the total the node shows. The JSON also lists the assumptions behind the figure. Costs nothing and needs no key - the numbers come from the replies the providers already sent. It sits at the end of the list rather than beside the other save_* switches because a widget inserted in the middle would shift every stored value after it in workflows saved before it existed.
reference_imagesoptIMAGEOptional look/identity references. Only the first 4 images of the batch are used, each downscaled to a 768 px longest side if it is larger; they are sent to the LLM to be described, and that description is forced into the art direction and the subject bible. The pixels themselves are prepended to the reference list of every start-frame render and of ref2va clips, so with 'image_provider' at 'pipe-steps' they still reach a ref2va video model while an i2va run keeps only the written description - and their presence is what switches the shots onto 'fal_image_edit_model'.

Outputs (6)

NameTypeDescription
pipeM2P_PIPEEverything this run wrote in words and numbers, on one wire: the start-frame and reference image prompts, the subject names, both MiniMax H3 prompt forms, the negatives, and every shot's index, start, end and duration, plus the transcript, the full analysis JSON and the per-shot audio clips. Feed it to 'Music2Video Pipe Expand' wherever you need one of them as its own socket; that node passes the pipe through, so it can be tapped as many times as you like along a chain. Only IMAGE and VIDEO stay on their own sockets here, because those go straight into a preview or a save node.
audio_clipsAUDIOOne AUDIO per shot, cut from the input track at that shot's boundaries and widened on both sides by 'audio_clip_padding', with the original sample rate and channel layout untouched. Always produced, even on a prompts-only run. Send them to PreviewAudio / SaveAudio or a lipsync node. The same clips are sent as driving audio when 'lipsync_audio' is on, 'video_provider' is fal, the chosen model declares an input for a driving audio track, and the clip itself lands inside that field's length window (2-15 s for a list-shaped reference field, 2-30 s otherwise). The clip is re-encoded to MP3 and sent inline in the request as a data: URI; clips outside the window go out without audio and a warning is logged.
imagesIMAGEThe rendered start frames, one IMAGE per shot in shot order. Empty when 'image_provider' is 'pipe-steps' (the default) - a prompts-only run puts nothing here, and any other setting means one billed image generation per shot. A shot whose render failed keeps its slot as a black frame so the list stays aligned with 'shot_index'; if every shot fails the node raises instead. Wire to PreviewImage / SaveImage, or into an image-to-video node as the first frame.
subject_imagesIMAGEThe rendered subject reference sheets, one per subject. Empty unless 'image_provider' is set AND 'render_subject_sheets' is on - each sheet is one more billed image on top of the per-shot renders. Sheets that failed are dropped rather than padded, so this can be shorter than 'reference_subjects'. These are the same images the run feeds back as identity references for the start frames, and - when 'video_prompt_source' is ref2va - as the <Picture N> references sent with the video request (the first nine references only, counting any images wired into 'reference_images').
videosVIDEOThe rendered per-shot clips as VIDEO, in shot order, opened from the mp4 files written to ComfyUI/output/music2prompts (or ComfyUI's temp folder when 'save_rendered_video' is off). Empty when 'video_provider' is 'pipe-steps' - any other setting means one billed video generation per shot - and empty on a ComfyUI too old to expose VideoFromFile. Clips that failed are skipped rather than padded, so this can be shorter than 'shot_index' - use 'final_video' if you need the cut in one piece.
final_videoVIDEOA single VIDEO: every rendered clip concatenated in shot order, each trimmed or its last frame held to its shot's exact duration, re-encoded through PyAV (libx264, yuv420p, mp4 with faststart). The output size is taken from the clips themselves - the most common frame size, ties going to the largest - and 'final_fit' decides whether an odd-sized clip is letterboxed, stretched or cropped into it. 'final_fps' sets the frame rate (0 = the highest rate found among the clips), 'final_crf' the x264 quality, and 'final_audio' the soundtrack. None unless 'video_provider' is set and 'concat_video' is on, None if the muxing pass fails (a warning is logged and the individual clips survive on 'videos'), and None on a ComfyUI too old to expose VideoFromFile - the mp4 is written to ComfyUI/output/music2prompts either way. Wire to SaveVideo.