H3 De-RoPE Stretch
Give fast motion more frames to live in
- vae
- images
- audio
- audio_vae
- latent
- model
- latent
- derope
- denoise
- report
- model
This is the most interesting node in the pack and probably the least understood. The pitch is: when a MiniMax H3 clip moves too fast, the fast frames come back smeared, because there simply are not enough latent rows to describe that much motion. H3 De-RoPE Stretch finds where the clip moves fastest, shows those frames several times over, encodes the stretched result with its audio, and hands you a latent to plug into a sampler. You sample it at a partial denoise - so the model redraws the smear using the extra room it now has - decode, and then H3 De-RoPE Recover puts the frames back on the clip's original timing.
Blur in, sharp out, same clip, same length. That is the loop.
How it works
H3 spaces its latent rows unevenly: one frame in the first row, then groups of four per row (the pack's own note spells out 1, 4, 4, 4, 4 frames per 17). So when a clip's motion spikes, several frames share too little representation, which is what you see as smear.
The node builds a motion profile per latent row, picks the rows above a threshold, and holds them - repeating a fast frame's latent row several times so the model gets more tokens to describe that motion in. Gaps between two held spans, up to a bridge width, get held too, so you do not get a stutter of sharp/smeared/sharp. Then it encodes the stretched clip plus its audio as your start latent.
mode picks the preset:
balanced- fastest quarter, held 4×wide- fastest 30%, held 4×economy- fastest 15%, held 3×manual- exposesthreshold,peak_hold,bridgeandrampif you want to tune it against a specific clip
Because this is a re-render, not an interpolation: the sampler keeps the clip's motion and redraws the smear.
The de-rope part
derope is a custom output type - a plan, not a picture. It carries what was held, in what order, with what peak, so H3 De-RoPE Recover can squeeze the extra frames back out afterwards. Lose that wire and you have an over-long clip: the frames are still there, several copies each, playing at a fraction of the original speed. The two nodes are a pair. Do not use one without the other.
Inputs and outputs
vae is the H3 video VAE. mode and audio_mode are dynamic combos.
strength is the number that sets your whole graph. It is passed out as denoise and is exactly what you feed the sampler: 0.5 keeps the clip's motion and redraws the smear, 0.7 redraws more and re-times the action. The author's note is specific - set the sampler's steps to the first pass's steps times this value, so 4 after an 8-step pass, 13 after 25.
audio_mode decides how much of the soundtrack the pass re-renders, and needs audio and audio_vae wired: follow 0.5, loose 0.7, pin keeps it untouched, fresh replaces it.
Everything else is optional: images (the clip's frames at 24 fps; leave empty to decode from latent), audio (from Get Video Components; empty, a joint H3 latent's own audio is used), audio_vae, fps, latent (a finished H3 latent - motion and audio read from it instead of re-encoding frames), model, and low_vram.
Outputs are latent (the stretched clip plus audio, for the sampler's latent_image), derope (for Recover), denoise (for the sampler), report (frames in and out, frames held, peak hold) and model.
latent has a sharp edge worth repeating: an early, unfinished estimate carries no motion to keep, so the profile finds nothing to hold and your clip comes back fast-forwarded. Feed it a finished latent, or the finished frames.
Install
ComfyUI Manager → WAS Node Suite v3 → install → restart, or:
cd ComfyUI/custom_nodes
git clone https://github.com/WASasquatch/was-node-suite-comfyui.git
ComfyUI 0.14.0+ and Python 3.10+. The pack installs no packages - requirements.txt is a comment, and nothing is fetched from the network while features.network is false in config.yaml. What you do need is the H3 video VAE and, if you want audio re-rendered, the audio VAE, both as ordinary ComfyUI models.
Two notes from the node's own description that are easy to miss: the GPU's models are released before and after its VAE work, so the samplers on either side of it get the whole card; and the strip on the node shows per-row motion and what was held, which is the fastest way to sanity-check that it is holding the right part of the clip.
Set low_vram on and it slices the model over the stretched clip at 8192 tokens per slice, streaming weights, for the same result at a lower peak. It needs model wired to do that.
Inputs (11)
| Name | Type | Default | Description |
|---|---|---|---|
| vae | VAE | The H3 video VAE. | |
| mode | COMBO | How much of the clip is held. `balanced` holds the fastest quarter at 4, `wide` the fastest 30% at 4, `economy` the fastest 15% at 3. `manual` shows every setting. | |
| strength | FLOAT | 0.500.05–1 | Denoise for the sampler, passed out as `denoise`: `0.5` keeps the clip's motion and redraws the smear, `0.7` redraws more and re-times the action. Set the sampler's steps to the first pass's times this, as `4` after an 8-step pass or `13` after 25. |
| audio_mode | COMBO | How much of the audio the pass re-renders: `follow` 0.5, `loose` 0.7, `pin` keeps it, `fresh` replaces it. Needs audio and audio_vae. | |
| imagesopt | IMAGE | The clip's frames at 24 fps. Leave empty to decode them from latent. | |
| audioopt | AUDIO | The clip's soundtrack, from Get Video Components. Left empty, a joint H3 latent's own audio is used. Without either, held spans come back rushed. | |
| audio_vaeopt | VAE | The H3 audio VAE, for encoding the audio and decoding a latent's. | |
| fpsopt | FLOAT | 24.0001–120 | Frame rate of the clip, as `24`, for timing its audio. |
| latentopt | LATENT | The clip's finished H3 latent, such as a sampler's output. Motion, audio and scene cuts are read from it instead of encoding the frames, and no hold crosses a cut. An early, unfinished estimate carries no motion to keep, and comes back fast-forwarded. | |
| modelopt | MODEL | The H3 model the sampler uses. Passed out for the sampler's model. | |
| low_vramopt | BOOLEAN | true | `true` runs each model block over the stretched clip in slices of 8192 tokens and keeps only the weights that fit beside it on the card, streaming the rest, for the same output at a lower memory peak. Needs model. |
Outputs (5)
| Name | Type | Description |
|---|---|---|
| latent | LATENT | The stretched clip and its audio, for the sampler's latent_image. A clip with cuts carries where each scene opens, for H3 Decode Video. |
| derope | WAS_H3_DEROPE | What was held, for H3 De-RoPE Recover. |
| denoise | FLOAT | The strength, for the sampler's denoise. |
| report | STRING | Frames in and out, frames held, the peak hold and the cuts kept out of the motion measure. |
| model | MODEL | The model for the sampler, run in slices when low_vram is on. |