Nodes/UIIIAIII Toolkit/Agnes Image to Video (agnes-video-v2.0)
ComfyUI Node

Agnes Image to Video (agnes-video-v2.0)

Animate a still, without owning the model

By uiiiaiii·Created 23 days ago·Updated 10 days ago· 0
Agnes Image to Video (agnes-video-v2.0)
  • image
  • video
◄prompt►
◄ratio16:9►
◄resolution720p►
◄duration5s►
◄frame_rate24►
◄negative_prompt►
◄seed0►
◄num_inference_steps0►
◄poll_interval5►
◄max_wait_time600►

Image-to-video is the job people actually want: take a portrait, a product shot, a piece of concept art, and give it a few seconds of believable motion. Doing that locally means a Wan-class checkpoint and a lot of VRAM patience. This node means you hand the picture to Agnes Video V2.0 and get a clip back.

It's the middle sibling of the pack's three Agnes video nodes: text-to-video writes from nothing, this one animates what you give it, and Keyframe Animation interpolates between two stills. Same API key, same settings panel, same async job underneath.

What's on the node

Required:

  • image - the still you want to move.
  • prompt - what should move, and how. The author's tooltip is explicit: "what should move." "Hair drifting, subtle head turn, camera slowly pushes in" beats "cinematic mood."
  • ratio - 16:9, 9:16, 1:1, 4:3, 3:4.
  • resolution - 480p, 720p, 1080p.
  • duration - 3s / 5s / 10s / 18s.
  • frame_rate - 1–60.

Optional: negative_prompt, seed (0 = unspecified; non-zero seeds are sent, so this node is actually reproducible), num_inference_steps (0 = API default), poll_interval (seconds between status checks), max_wait_time (default 600s, up to an hour).

Output: video (VIDEO). On ComfyUI builds without the video type, it falls back to a video_path string pointing at an .mp4 in your output directory.

The gotcha worth knowing before you tune anything

For image-driven modes the pack does not send width/height to the API - the upstream spec derives output dimensions from the input image. So ratio and resolution are effectively inert here: you can set 9:16 on a landscape photo and you'll get a landscape clip, because the size came from your image. Those two dropdowns decide the frame for text-to-video; for image-to-video, your source image decides.

Practically: crop or pad your input to the shape you want before this node. That's a much faster feedback loop than re-queuing a two-minute job and discovering the frame is wrong. It also means a huge input image produces a huge output clip - downscale to roughly your target first if you care about turnaround.

How it runs

The node encodes your image to a PNG data URI, POSTs a video task, receives a video_id, then polls with a progress callback that prints status into the node until the job completes, then downloads the file. ComfyUI's queue is blocked while that happens - nothing downstream of this node executes during the wait, and the node is an output node, so there's no trick of moving it off the critical path.

One advisory straight from the pack's own source comments: Agnes' spec expects a public URL for the input image, and since a local ComfyUI has no way to serve one, the pack sends a data URI as the fallback and notes that if the API doesn't accept that, you'd need to host the image yourself. In practice the data URI path is what everything uses - but if you see image-to-video failing while text-to-video works, that's the mechanism to suspect, not your prompt.

Frame counts come from the duration preset (81 / 121 / 241 / 441 frames, all valid under the 8n+1 rule), and frame_rate is sent separately. Crank frame_rate to 60 and a "5s" preset becomes a two-second clip at 24fps-equivalent length, because the frame count is the fixed quantity.

Install

cd ComfyUI/custom_nodes
git clone https://github.com/uiiiaiii/UIIIAIII_Toolkit.git
# restart ComfyUI

Manager users: search UIIIAIII Toolkit. The only declared dependency is requests>=2.28.0; there is no model download and no ffmpeg requirement, because the API hands back a finished MP4.

Then key it. platform.agnes-ai.com for the key, Settings → UIIIAIII Toolkit → ① Node API to store it. That writes config.json inside the pack directory; the config file takes precedence over an AGNES_API_KEY environment variable for these nodes, and the key sits there in plaintext. If you wouldn't paste a credential into a random repo's config file, read this one first - it's short, and you should be suspicious of any pack whose stated purpose is sending your data and your key somewhere.

Common problems

"Input image encoding failed." Wrong socket contents - a MASK or a malformed tensor. Wire a real IMAGE.

Polling timed out. Long clips on a busy free tier. Raise max_wait_time; it goes up to 3600 seconds.

The clip barely moves, or moves too much. That's prompt craft, not settings. Name the motion and its magnitude, and keep the subject description short - the model already has the picture, so the prompt is a direction, not a description.

Output length feels short. Check frame_rate. The duration preset fixes the frame count, so a non-24 frame rate rescales the wall-clock length.

ComfyUI 0.29 or older. You'll get a file path instead of a VIDEO object. The clip is fine; it's just saved to disk and referenced by string, so downstream video nodes expecting VIDEO won't accept it.

CategoryUIIIAIII Toolkit/Agnes

Inputs (11)

NameTypeDefaultDescription
promptSTRINGText description of the video content (what should move)
imageIMAGEInput image (uploaded to the Agnes API)
ratioCOMBO16:9Video aspect ratio
resolutionCOMBO720pVideo resolution preset
durationCOMBO5sVideo duration (based on num_frames and frame_rate)
frame_rateINT241–60Video frame rate (1-60)
negative_promptoptSTRINGNegative prompt describing what to avoid
seedoptINT00–18446744073709550000Random seed (0 means not specified)
num_inference_stepsoptINT00–200Inference steps (0 uses the default value)
poll_intervaloptINT51–60Task polling interval (seconds)
max_wait_timeoptINT60060–3600Maximum wait time (seconds)

Outputs (1)

NameTypeDescription
videoVIDEO—