Nodes/ComfyUI-Kling-Direct/Kling 3.0 Turbo Text to Video
ComfyUI Node

Kling 3.0 Turbo Text to Video

Storyboards in a Textbox

By IxMxAMAR·Created 6 months ago·Updated 3 days ago· 5
Kling 3.0 Turbo Text to Video
  • auth
  • video
  • video_file
  • audio
  • url
  • task_id
◄prompt►
◄resolution720p►
◄aspect_ratio16:9►
◄duration5►

No image needed. Kling 3.0 Turbo Text to Video takes a prompt, a resolution, an aspect ratio and a duration, and returns a clip with sound. It's the fastest way in this pack to go from words to a moving picture, and the input list is small enough that the interesting part is what you put in the prompt - which has its own tiny syntax.

The prompt has a storyboard format

The tooltip is the documentation here, and it's worth reading twice:

shot 1, 3, a wolf pads through fresh snow; shot 2, 2, it lifts its head and howls;

Shot number, seconds, prompt. That's a multi-shot storyboard written in a text field - multiple cuts in one generation, with per-shot durations, instead of you wrangling three separate jobs and hoping they match. For something that sounds like a marketing bullet point, it's genuinely the most interesting thing about this node: short-form video work is mostly pacing, and pacing is the thing you can now describe directly.

Prompts also get normalized before they're sent. The pack converts @image1 / @video1 style references into <<<image_1>>> tokens and enforces a 2500-character ceiling, so a runaway prompt gets rejected with a clear length error rather than silently truncated. Keep storyboard prompts tight - three shots in 2500 characters is easy, but not if each shot is a paragraph.

Inputs and outputs

  • prompt - the text, multiline, storyboard-capable.
  • resolution - 720p or 1080p.
  • aspect_ratio - 16:9, 9:16, 1:1. Use 9:16 deliberately; vertical is where a lot of this model's output actually gets used, and generating landscape and cropping is a waste of a call.
  • duration - an integer from 3 to 15 seconds. This is the total, so make your shot seconds add up to it or the model decides the rest.

Outputs: video (IMAGE frame batch), video_file (filename in your output directory), audio (the clip's audio track - this node exposes no sound toggle, so you take what comes back), url (Kling's CDN link, which expires), and task_id.

The API key thing

Turbo is a distinct API surface, and it authenticates differently from everything else in this pack. Set the api_key field on Kling AI Authentication - or export KLING_API_KEY - and wire the auth output in. The older access-key/secret-key pair is not enough, and the node tells you so rather than failing obscurely. Requests go to /text-to-video/kling-3.0-turbo on the unified /tasks API, which means Task Status can poll the same job if your workflow got interrupted.

Get the key from https://kling.ai/dev. New accounts need KYC activated first, which is the actual friction in this whole exercise - not the install.

Install

ComfyUI Manager → Install Custom Nodes → search "Kling Direct" → Install → restart. Or:

cd ComfyUI/custom_nodes
git clone https://github.com/IxMxAMAR/ComfyUI-Kling-Direct

The pack has no model downloads and no heavy dependencies - stdlib plus requests, Pillow, numpy, torch and opencv-python, all already present in a working ComfyUI. Your "install" is really account setup, and the pack ships a Kling API Health Check node that pings the free /account/costs endpoint so you can confirm auth before you spend anything. If you're outside the Singapore region, insert Kling Region Selector after Auth: china and us are options, and a wrong region fails in a way that looks like a key problem.

What to expect once you press queue

Text to video is the slowest thing in this pack because there's no reference image to anchor on - the model is generating the whole scene. The node blocks and polls, with intervals backing off to 15 seconds after two minutes and 30 after five, and a 1200-second ceiling before it raises a timeout. Cancelling mid-poll propagates within about a second, so you're not trapped; you just don't get the credits back.

Expect quality to be iterative, not first-try. Closed video models reward specific, physical prompts - subject, action, camera, light - and punish adjective piles the same way local models do. Start at 720p and 5 seconds, get a take you like, then spend the 1080p money. And check Cost Estimator before queueing a batch of storyboard variants, because the per-call pricing on this tier adds up fast enough that nobody wants to discover it from their card statement.

CategoryKling AI/Video

Inputs (5)

NameTypeDefaultDescription
authKLING_AUTH—
promptSTRINGText description of the video. Multi-shot format: 'shot 1, 3, words; shot 2, 2, words;' (shot number, seconds, prompt).
resolutionCOMBO720pOutput resolution.
aspect_ratioCOMBO16:9Output video aspect ratio.
durationINT53–15Video duration in seconds.

Outputs (5)

NameTypeDescription
videoIMAGE—
video_fileSTRING—
audioAUDIO—
urlSTRING—
task_idSTRING—