Nodes/was-node-suite-comfyui/H3 De-RoPE Stretch
ComfyUI Node Runs on cloud

H3 De-RoPE Stretch

Give fast motion more frames to live in

By WASasquatch·Created 4 years ago·Updated a day ago· 1,864
H3 De-RoPE Stretch
  • vae
  • images
  • audio
  • audio_vae
  • latent
  • model
  • latent
  • derope
  • denoise
  • report
  • model
◄mode▾►
◄strength0.50►
◄audio_mode▾►
◄fps24.000►
◄low_vramtrue►

This is the most interesting node in the pack and probably the least understood. The pitch is: when a MiniMax H3 clip moves too fast, the fast frames come back smeared, because there simply are not enough latent rows to describe that much motion. H3 De-RoPE Stretch finds where the clip moves fastest, shows those frames several times over, encodes the stretched result with its audio, and hands you a latent to plug into a sampler. You sample it at a partial denoise - so the model redraws the smear using the extra room it now has - decode, and then H3 De-RoPE Recover puts the frames back on the clip's original timing.

Blur in, sharp out, same clip, same length. That is the loop.

How it works

H3 spaces its latent rows unevenly: one frame in the first row, then groups of four per row (the pack's own note spells out 1, 4, 4, 4, 4 frames per 17). So when a clip's motion spikes, several frames share too little representation, which is what you see as smear.

The node builds a motion profile per latent row, picks the rows above a threshold, and holds them - repeating a fast frame's latent row several times so the model gets more tokens to describe that motion in. Gaps between two held spans, up to a bridge width, get held too, so you do not get a stutter of sharp/smeared/sharp. Then it encodes the stretched clip plus its audio as your start latent.

mode picks the preset:

  • balanced - fastest quarter, held 4×
  • wide - fastest 30%, held 4×
  • economy - fastest 15%, held 3×
  • manual - exposes threshold, peak_hold, bridge and ramp if you want to tune it against a specific clip

Because this is a re-render, not an interpolation: the sampler keeps the clip's motion and redraws the smear.

The de-rope part

derope is a custom output type - a plan, not a picture. It carries what was held, in what order, with what peak, so H3 De-RoPE Recover can squeeze the extra frames back out afterwards. Lose that wire and you have an over-long clip: the frames are still there, several copies each, playing at a fraction of the original speed. The two nodes are a pair. Do not use one without the other.

Inputs and outputs

vae is the H3 video VAE. mode and audio_mode are dynamic combos.

strength is the number that sets your whole graph. It is passed out as denoise and is exactly what you feed the sampler: 0.5 keeps the clip's motion and redraws the smear, 0.7 redraws more and re-times the action. The author's note is specific - set the sampler's steps to the first pass's steps times this value, so 4 after an 8-step pass, 13 after 25.

audio_mode decides how much of the soundtrack the pass re-renders, and needs audio and audio_vae wired: follow 0.5, loose 0.7, pin keeps it untouched, fresh replaces it.

Everything else is optional: images (the clip's frames at 24 fps; leave empty to decode from latent), audio (from Get Video Components; empty, a joint H3 latent's own audio is used), audio_vae, fps, latent (a finished H3 latent - motion and audio read from it instead of re-encoding frames), model, and low_vram.

Outputs are latent (the stretched clip plus audio, for the sampler's latent_image), derope (for Recover), denoise (for the sampler), report (frames in and out, frames held, peak hold) and model.

latent has a sharp edge worth repeating: an early, unfinished estimate carries no motion to keep, so the profile finds nothing to hold and your clip comes back fast-forwarded. Feed it a finished latent, or the finished frames.

Install

ComfyUI Manager → WAS Node Suite v3 → install → restart, or:

cd ComfyUI/custom_nodes
git clone https://github.com/WASasquatch/was-node-suite-comfyui.git

ComfyUI 0.14.0+ and Python 3.10+. The pack installs no packages - requirements.txt is a comment, and nothing is fetched from the network while features.network is false in config.yaml. What you do need is the H3 video VAE and, if you want audio re-rendered, the audio VAE, both as ordinary ComfyUI models.

Two notes from the node's own description that are easy to miss: the GPU's models are released before and after its VAE work, so the samplers on either side of it get the whole card; and the strip on the node shows per-row motion and what was held, which is the fastest way to sanity-check that it is holding the right part of the clip.

Set low_vram on and it slices the model over the stretched clip at 8192 tokens per slice, streaming weights, for the same result at a lower peak. It needs model wired to do that.

CategoryWAS Suite/Latent/Video

Inputs (11)

NameTypeDefaultDescription
vaeVAEThe H3 video VAE.
modeCOMBOHow much of the clip is held. `balanced` holds the fastest quarter at 4, `wide` the fastest 30% at 4, `economy` the fastest 15% at 3. `manual` shows every setting.
strengthFLOAT0.500.05–1Denoise for the sampler, passed out as `denoise`: `0.5` keeps the clip's motion and redraws the smear, `0.7` redraws more and re-times the action. Set the sampler's steps to the first pass's times this, as `4` after an 8-step pass or `13` after 25.
audio_modeCOMBOHow much of the audio the pass re-renders: `follow` 0.5, `loose` 0.7, `pin` keeps it, `fresh` replaces it. Needs audio and audio_vae.
imagesoptIMAGEThe clip's frames at 24 fps. Leave empty to decode them from latent.
audiooptAUDIOThe clip's soundtrack, from Get Video Components. Left empty, a joint H3 latent's own audio is used. Without either, held spans come back rushed.
audio_vaeoptVAEThe H3 audio VAE, for encoding the audio and decoding a latent's.
fpsoptFLOAT24.0001–120Frame rate of the clip, as `24`, for timing its audio.
latentoptLATENTThe clip's finished H3 latent, such as a sampler's output. Motion, audio and scene cuts are read from it instead of encoding the frames, and no hold crosses a cut. An early, unfinished estimate carries no motion to keep, and comes back fast-forwarded.
modeloptMODELThe H3 model the sampler uses. Passed out for the sampler's model.
low_vramoptBOOLEANtrue`true` runs each model block over the stretched clip in slices of 8192 tokens and keeps only the weights that fit beside it on the card, streaming the rest, for the same output at a lower memory peak. Needs model.

Outputs (5)

NameTypeDescription
latentLATENTThe stretched clip and its audio, for the sampler's latent_image. A clip with cuts carries where each scene opens, for H3 Decode Video.
deropeWAS_H3_DEROPEWhat was held, for H3 De-RoPE Recover.
denoiseFLOATThe strength, for the sampler's denoise.
reportSTRINGFrames in and out, frames held, the peak hold and the cuts kept out of the motion measure.
modelMODELThe model for the sampler, run in slices when low_vram is on.