FL_MiniMaxH3MotionPrepare
The Half of Motion Refine That Only Runs Once
- vae
- positive
- latent
- audio_vae
- baseline_images
- baseline_audio
- shot_data
- FL_H3_MOTION_PREP
You will almost never place this node by hand. FL MiniMax H3 Motion Prepare is one of the two internal stages that FL MiniMax H3 Motion Refine expands into at run time - it's flagged dev-only, so it doesn't even show up in the node menu unless you've turned ComfyUI's dev mode on. It's worth understanding anyway, because it owns the expensive half of that pipeline and it's where nearly every Motion Refine error message actually comes from.
What it's for
Motion Refine does two very different jobs: prepare (decode the baseline, stretch the fast motion, resample audio, re-encode everything) and sample (partial denoise, recover original timing). Splitting them is what makes iteration bearable. Preparation is cached by ComfyUI as its own node, so once it's done you can change the seed, the step count, the strength, the sampler, the scheduler or the audio strength and reuse it. Change the source, the VAEs, motion coverage or the resize inputs and it's thrown away. That's the whole reason the parent node is built this way.
Inputs that matter
Required: vae, target_long_side (0 keeps the source size; otherwise the long edge is upscaled, aligned to 32-pixel H3 canvases), motion_coverage, max_hold (2–8, only consulted by uniform - the adaptive presets own their own hold limits), expand_to_end (extends a short unheld tail through the last motion burst, matching H3 Time Smear), and upscale_method.
Optional: positive - mandatory unless you're in shot-plan mode, where each shot's own conditioning is used instead - plus latent, audio_vae, baseline_images, baseline_audio, and shot_data. Its single output is an FL_H3_MOTION_PREP bundle: stretched video latent, stretched audio latent, the hold map, retimed conditioning, source and output dimensions, and a timing string.
What it does inside
It first refuses anything that isn't a completed native H3 clip. The video latent has to use H3's 5n+2 latent positions, from which it derives the frame count, and the audio latent has to match that duration - a mismatch is a hard error rather than a silent trim.
Then it builds the hold map. For balanced, economical and wide it asks MAINodes' H3JerkOracle how jerky each interval is and gets a preset-shaped map back. uniform uses your max_hold everywhere. off gives every frame a hold of 1, which means no temporal stretch at all - that's your spatial-only benchmark path, and it's also the only path where audio_vae is optional.
H3TimeSmear.plan turns the map into actual hold counts and an expanded frame total on H3's grid (a 39-frame internal floor). In shot-plan mode it also protects the hidden motion-context prefix and the tail from stretching, so the previous shot's conditioning isn't blown apart by the expansion.
The decoding and re-encoding come next: decode the baseline (or take your baseline_images), upscale the pixels if you asked it to, then re-encode with an index list that repeats each frame according to its hold - one stretch, in VAE space with the original temporal arithmetic intact. Audio gets the same treatment via H3AudioSmear, then re-encodes through the audio VAE. If the timeline didn't expand, the source audio latent is passed straight through.
Finally it retimes conditioning: image anchors are resized to the new canvas, temporal masks get re-quantized onto the stretched clock, and prior-shot context is applied when there is any.
Install, and the errors you'll actually see
Same as the rest of the pack - ComfyUI Manager, search FL MiniMax H3, or:
cd ComfyUI/custom_nodes
git clone https://github.com/filliptm/ComfyUI-FL-MiniMaxH3.git
Restart, and install/update ComfyUI-MAINodes; preparation is where the oracle and the smear operations get called, so a missing or outdated MAINodes fails here first. Current ComfyUI is required too - the pack imports ComfyUI's native H3 module.
The messages worth knowing:
source video must use H3's 5n+2 latent positions- you're feeding a clip that didn't come from H3's own grid.adaptive coverage needs at least 22 source frames- short renders can't be jerk-analysed; useuniformoroff.connect the H3 audio VAE to stretch the baseline audio- you picked a stretching coverage mode without anaudio_vae.shared baseline images must match the source latent's frame count and dimensions- yourbaseline_imagescame from a different latent, or from an already-refined output. In shot-plan mode they must be the full assembled timeline, not one shot.
One practical note: the timing string on the prep output describes the original preparation. If you're looking at cached timings, you're reading history, not the current render.
Inputs (12)
| Name | Type | Default | Description |
|---|---|---|---|
| vae | VAE | — | |
| target_long_side | INT | 00–16384 | 0 keeps source size. Otherwise upscale the long edge; dimensions align to 32 pixels. |
| motion_coverage | COMBO | balanced | Off benchmarks spatial refinement without temporal stretching. Uniform uses max_hold everywhere. |
| max_hold | INT | 42–8 | Used only by uniform coverage. Adaptive coverage presets own their hold limits. |
| expand_to_end | BOOLEAN | true | Extend a short unheld tail through the last motion burst, matching H3 Time Smear. |
| upscale_method | COMBO | lanczos | 3 options: lanczos, bicubic, bilinear |
| positiveopt | CONDITIONING | Required without shot_plan. Shot-plan mode uses each shot's own conditioning instead. | |
| latentopt | LATENT | Completed native H3 video/audio latent, not an empty latent or temporal reshot. | |
| audio_vaeopt | VAE | Required when the timeline expands; not needed with coverage off. | |
| baseline_imagesopt | IMAGE | Optional shared decode of this exact source latent. Connect the same images used by the before panel. | |
| baseline_audioopt | AUDIO | Optional shared decode of this exact source latent's soundtrack. | |
| shot_dataopt | FL_H3_MOTION_SHOT | — |
Outputs (1)
| Name | Type | Description |
|---|---|---|
| FL_H3_MOTION_PREP | FL_H3_MOTION_PREP | — |