Bonfire Shot
Five Seconds at 24 FPS Is Not 120 Frames, and H3 Knows It
- image
- width
- height
- duration
You want a five-second clip at 24 fps, so you want 120 frames. Not if you're running MiniMax H3. H3 only accepts frame counts where frames % 17 == 5, and 120 isn't one. The next legal count is 124, which at 24 fps is 5.17 seconds - slightly more clip than you asked for, and the kind of half-second surprise you discover at the end of a run.
Bonfire Shot does that arithmetic for you: a grid-aligned target resolution plus a legal frame count, without touching the connected image.
What it's for
Video graphs make you think in three units at once - pixels, seconds, frames - and every model has its own idea of which values are legal. Wan's VAE compresses time about 4x, which is why the 4n+1 rule and "why did my clip come back two frames short" are permanent fixtures of Wan threads. H3 is the same problem with a different modulus: 17, remainder 5, over a trained 124–362 frame range - the source of its "4 to 15 seconds" claim.
So this node answers what am I asking the model to make as one decision, before anything expensive happens. It's arithmetic, not a sampler, and it's the same species as the primitive value nodes: a value defined once instead of retyped into a latent node's width, height and length boxes.
How the frame count actually comes out
Seconds times fps, rounded half-up (not Python's banker's rounding, which would send 2.5 to 2). Then, on the H3 profile, snapped up to the next count where frames % 17 == 5, capped at 3600 frames. So 5s at 24 gives 120 → 124. A duration outside H3's trained 124–362 window still runs, with a warning that quality may drop, and the node shows both the requested duration and the real one. The frame count is what the H3 latent node wants in its length box.
Resolution works the same way: pick a mode, pick an alignment grid, both axes land on that grid. alignment is the same set the loader uses - MiniMax H3 (32), Krea 2 (16), plain multiples of 8/32/64, or a custom_multiple. Model profiles snap; bare multiples just report what you typed.
Three modes decide where the numbers come from. aspect_mp takes an aspect_ratio preset and a megapixels target and solves w × h = pixels with w / h = ratio. custom_size uses custom_width and custom_height as-is. match_input reads dimensions from the optional width / height sockets or the optional image - read for its shape, never modified or returned, so nothing downstream sees an altered picture. A socket wins on its axis, the image fills in the rest, and a missing axis gets a validation message rather than a guess.
One expectation to calibrate: snapping means your requested ratio isn't exact. 16:9 at 1 megapixel on the 32 grid comes out 1344×736 - about 0.99MP. The node prints the real numbers instead of pretending.
The fields that matter
Set resolution_mode and alignment, then either an aspect_ratio + megapixels pair or the two custom sizes. For length, seconds and fps - and fps is arithmetic only. It converts seconds into frames; it doesn't set the playback rate of what you encode later.
The outputs are width, height and duration. Wire the first two into your video latent node's width and height, duration into its length box. The name is the trap: duration is a frame count, not seconds. It's the one thing the README stops to say out loud, which is usually a sign the author watched people get it wrong.
Install
Manager → search Bonfire, or:
cd ComfyUI/custom_nodes
git clone https://github.com/DaemonSK/ComfyUI-Bonfire ComfyUI-Bonfire
Restart. No pip install, no model downloads - the pack declares no dependencies and this node is pure Python around the graph. It ships in the same two-node pack as Bonfire Image Loader, so wanting one gets you both. Same version caveat: it's a V3-API pack, so run a current ComfyUI.
Gotchas
Typing seconds into the length box. Give your latent node 5 and you've asked for 5 frames, about a fifth of a second. Use the duration output.
Expecting exactly the ratio or megapixels you typed. Snapping moves both axes to the nearest legal multiple; read the line at the bottom of the node.
Assuming match_input with nothing connected does nothing. It raises a validation error, which beats inventing 1024×1024 for you.
Using the H3 profile for a model that isn't H3. The default is shaped around H3's grid. On Wan, LTX or anything else, pick the matching grid or a plain multiple.
And one non-technical note, since the defaults here are so H3-shaped: H3's weights ship under a community licence that excludes the US, EU, UK and South Korea, outputs included - so for some readers the default profile describes a model they aren't licensed to run locally. The arithmetic holds for whatever video model you are running.
Inputs (12)
| Name | Type | Default | Description |
|---|---|---|---|
| resolution_mode | COMBO | aspect_mp | Where the output dimensions come from. |
| alignment | COMBO | minimax_h3 | Pixel grid both output dimensions are rounded to. |
| custom_multiple | INT | 81–1024 | Grid used when the alignment profile is 'custom'. |
| aspect_ratio | COMBO | 16:9 | Ratio used by the 'aspect_mp' resolution mode. |
| megapixels | FLOAT | 1.000.01–64 | Target area used by the 'aspect_mp' resolution mode. |
| custom_width | INT | 10241–16384 | Width used by the 'custom_size' resolution mode. |
| custom_height | INT | 10241–16384 | Height used by the 'custom_size' resolution mode. |
| seconds | FLOAT | 5.00.1–150 | Requested duration in seconds, before the frame grid. |
| fps | INT | 241–240 | Frames per second used to convert seconds to frames. |
| widthopt | INT | Match input: overrides the width taken from the image. | |
| heightopt | INT | Match input: overrides the height taken from the image. | |
| imageopt | IMAGE | Match input: read for its dimensions only. Never modified and never returned. |
Outputs (3)
| Name | Type | Description |
|---|---|---|
| width | INT | — |
| height | INT | — |
| duration | INT | — |