Nodes/ComfyUI-Kling-Direct/Kling 3.0 Turbo Image to Video
ComfyUI Node

Kling 3.0 Turbo Image to Video

The Fast Path Costs an API Key

By IxMxAMAR·Created 6 months ago·Updated 3 days ago· 5
Kling 3.0 Turbo Image to Video
  • auth
  • image
  • video
  • video_file
  • audio
  • url
  • task_id
◄prompt►
◄resolution720p►
◄duration5►

Kling 3.0 landed in February 2026 and, as usual, there was no download link - it's Kuaishou's closed model, so the only door is the API. Kling 3.0 Turbo Image to Video is the pack's handle on the fast tier of it: hand it a still, optionally a prompt, pick 720p or 1080p, pick a duration from 3 to 15 seconds, and it comes back with a clip, plus an audio output carrying whatever the returned file has in it.

If you're the "not through an API, go kick rocks" type, this isn't for you and the objection is fair - you're metering every take and your source image leaves your machine. But if you've wanted a closed video model inside a graph next to your own upscaler and your own masking, this is how you get it, and the fact that the whole thing is one node with four inputs is the whole appeal.

What's different about Turbo

Turbo is not a parameter on the standard image-to-video node - it's a separate code path hitting a separate API surface. The pack routes it to /image-to-video/kling-3.0-turbo with a contents array (an optional prompt entry, then a first_frame entry carrying your base64 image) and a settings block with resolution and duration. The older video endpoints use per-model request shapes; Turbo uses a unified one that ComfyUI's Task Status node can also query through /tasks.

The catch, and it's the thing to remember: Turbo requires an API key. Not the access-key/secret-key pair you use everywhere else in this pack - a single API key set on the Kling AI Authentication node (or in the KLING_API_KEY environment variable). Wire it without one and the node fails immediately with a message telling you exactly that, which is merciful.

The inputs

Four, and there's nothing hiding in them.

  • image - your first frame. This one conditions the animation directly rather than acting as a loose subject reference, so unlike the pack's multi-image nodes, you're not handing it a cropped headshot; you're handing it the shot.
  • prompt - optional, and the tooltip says so. "Optional" is doing real work: a still with clear motion cues sometimes needs no text at all, and a prompt that fights the image is worse than an empty field.
  • resolution - 720p or 1080p. 720p for iteration, 1080p once the take is right. Nobody's evaluation is improved by iterating at 1080p.
  • duration - an integer, 3 to 15 seconds.

Outputs are the pack's standard video quintet: video (frame batch), video_file (the .mp4 written into ComfyUI's output dir), audio, url (expiring), and task_id.

Install

ComfyUI Manager → Install Custom Nodes → search "Kling Direct" → Install → restart. Manual:

cd ComfyUI/custom_nodes
git clone https://github.com/IxMxAMAR/ComfyUI-Kling-Direct

No weights, no models folder, no torch upgrade - the pack depends only on stdlib plus requests, Pillow, numpy, torch and opencv-python, which ComfyUI already has. The real setup is on Kling's side: an account at https://kling.ai/dev, KYC activation, then an API key. While you're there, use the pack's Kling API Health Check node once - it calls the free /account/costs endpoint and tells you whether auth actually works, which saves a lot of staring at a spinner.

Where this bites

The node polls, so your workflow blocks until the job is done. Polling intervals stretch to 15 seconds past two minutes and 30 past five, with a 1200-second timeout - on a 15-second 1080p Turbo job that's a long, quiet canvas. Cancelling does propagate to ComfyUI within about a second, and the progress bar is an estimate, not a real percentage; Kling doesn't expose one, and the changelog admits as much.

Region is the other classic. The pack defaults to Singapore; accounts in mainland China need the Kling Region Selector set to china, and US accounts to us. Wrong region looks exactly like bad credentials, which is why it wastes people's afternoons.

And do the arithmetic before you queue a batch. Closed video is priced per call - the same tier people quote at around $1.50 for a five-second 1080p clip on Kling's Omni models - so a fifteen-second 1080p Turbo run is not a "let's see what happens" experiment. That's what Cost Estimator is for.

CategoryKling AI/Video

Inputs (5)

NameTypeDefaultDescription
authKLING_AUTH—
imageIMAGE—
promptSTRINGOptional text prompt to guide the video generation from the image.
resolutionCOMBO720pOutput resolution.
durationINT53–15Video duration in seconds.

Outputs (5)

NameTypeDescription
videoIMAGE—
video_fileSTRING—
audioAUDIO—
urlSTRING—
task_idSTRING—