Academia SD Moviola Out
Keep the motion, not just the last PNG
- images
- latent
- path
- latent_frames
Moviola Out is the tail of the chain: it saves the last frame of the take you just generated, plus a piece of the latent that produced it, so the next pass has something better than a JPEG to continue from. It's also the node with the setting people don't understand and then mis-tune - latent_frames - so let's start there.
Why a latent frame beats a PNG
Here's the thing that makes last-frame chaining look worse than it should. Encode a final frame as a PNG, decode it back and feed it in as the first frame of the next clip, and you've handed the model a photograph. It has no idea whether the camera was panning, whether the subject was mid-turn, whether anything was moving at all.
H3's video VAE compresses time in blocks of four - the pack spells out the ratio as FRAME_PER_TOKEN = (1, 4, 4, 4, 4), meaning every latent frame after the first encodes four real frames. So the final latent frame isn't a still; it carries the direction and speed of the motion inside it. Anchor from that and the next take picks up the movement. Anchor from a PNG and the camera can quite happily start off travelling the other way, which is the classic "why does take six open with a jerk" symptom.
Moviola Out saves both: toma_00002_.png with the last frame, and toma_00002_.safetensors with the latent, numbered as a pair. Moviola In reads the PNG, Moviola Guide reads the latent.
The inputs and outputs
project_path- the take identity, same string as In and Guide (defaultmoviola/toma). It's resolved under ComfyUI'soutputfolder, and a path with..in it raises rather than writing outside.images(optional) - the decoded video. Only the last frame gets written; keeping every frame would fill your drive for nothing, and you have a video combine node for the full clip.latent(optional) - the sampler's output latent. This is what becomes the.safetensorshandoff.latent_frames- how many trailing latent frames to keep, 1 to 8. At1, you save exactly the motion token described above. Higher values keep more trajectory, at a real cost: each extra latent frame is roughly four real frames that the next take will open by replaying, so you trim it back off in the edit. Start at1. Only go higher if your shots keep losing momentum between takes - and then only to 2.path(STRING, output) - the base name it just wrote, without extension (.../output/moviola/toma_00002_). Handy as a filename prefix so your renders are stamped with which take produced them.
Connect at least one of images or latent, or it raises instead of silently doing nothing. The node is also flagged as an output node, so a workflow that ends here will actually run even if nothing downstream consumes path.
How the files get written
It finds the highest-numbered frame in the folder, adds one, and writes the PNG. The latent goes to a temporary file first and is then moved into place with an atomic rename - that's not paranoia, it's so a half-written file can never be sitting there for the next pass to read. The latent itself is handled shape-agnostically: H3 packs video and audio into a nested tensor and only the video half matters here, while a plain image latent is already a single frame.
Install
cd ComfyUI/custom_nodes
git clone https://github.com/AcademiaSD/comfyui_AcademiaSD
# restart ComfyUI, then Add Node → Academia SD/Moviola
Nothing extra to install - Pillow, torch and safetensors are all ComfyUI's own. The usual caveat with this trio: the pack README documents everything else in the collection but not Moviola, which arrived in the most recent commit, and there's no example workflow or community discussion for it yet.
Where people get burned
One project_path per loop. In and Out share the same folder scheme, so if you run two loops through the same take name you'll interleave their numbering and neither one will chain cleanly. Give each pipeline its own name.
Numbers are the counter, not the filenames you like. Out always writes highest-plus-one. If you delete toma_00003_ because you didn't like take three, the numbering carries on and In will keep serving the newest frame on disk - the folder, not your intent, decides.
Files that don't pair. Guide needs the .safetensors that matches the highest PNG. If you clean up "just the big latent files" to save space, Guide quietly stops anchoring and your cuts start showing. Keep them together or clear the folder entirely.
Changing resolution mid-loop. The saved latent belongs to a specific geometry. Change the clip size and the next take can't use it - connect the clip's latent to Moviola Guide's av_latent so you get a readable error instead of a deep traceback that never mentions Moviola.
This node saves the handoff, not your movie. The output PNG is a work file for the next take. Your actual video still needs a proper save / video combine node, or you'll finish a six-take shoot with six single frames and a lot of regret.
Inputs (5)
| Name | Type | Default | Description |
|---|---|---|---|
| project_path | STRING | moviola/toma | — |
| latent_frames | INT | 11–8 | — |
| frames_back | INT | -1-1–32 | -1 derives it from latent_frames, which is what you want. 0 keeps the last frame. |
| imagesopt | IMAGE | — | |
| latentopt | LATENT | — |
Outputs (2)
| Name | Type | Description |
|---|---|---|
| path | STRING | — |
| latent_frames | INT | — |