Kling 3.0 Turbo Text to Video
Storyboards in a Textbox
- auth
- video
- video_file
- audio
- url
- task_id
No image needed. Kling 3.0 Turbo Text to Video takes a prompt, a resolution, an aspect ratio and a duration, and returns a clip with sound. It's the fastest way in this pack to go from words to a moving picture, and the input list is small enough that the interesting part is what you put in the prompt - which has its own tiny syntax.
The prompt has a storyboard format
The tooltip is the documentation here, and it's worth reading twice:
shot 1, 3, a wolf pads through fresh snow; shot 2, 2, it lifts its head and howls;
Shot number, seconds, prompt. That's a multi-shot storyboard written in a text field - multiple cuts in one generation, with per-shot durations, instead of you wrangling three separate jobs and hoping they match. For something that sounds like a marketing bullet point, it's genuinely the most interesting thing about this node: short-form video work is mostly pacing, and pacing is the thing you can now describe directly.
Prompts also get normalized before they're sent. The pack converts @image1 / @video1 style references into <<<image_1>>> tokens and enforces a 2500-character ceiling, so a runaway prompt gets rejected with a clear length error rather than silently truncated. Keep storyboard prompts tight - three shots in 2500 characters is easy, but not if each shot is a paragraph.
Inputs and outputs
prompt- the text, multiline, storyboard-capable.resolution-720por1080p.aspect_ratio-16:9,9:16,1:1. Use9:16deliberately; vertical is where a lot of this model's output actually gets used, and generating landscape and cropping is a waste of a call.duration- an integer from 3 to 15 seconds. This is the total, so make your shot seconds add up to it or the model decides the rest.
Outputs: video (IMAGE frame batch), video_file (filename in your output directory), audio (the clip's audio track - this node exposes no sound toggle, so you take what comes back), url (Kling's CDN link, which expires), and task_id.
The API key thing
Turbo is a distinct API surface, and it authenticates differently from everything else in this pack. Set the api_key field on Kling AI Authentication - or export KLING_API_KEY - and wire the auth output in. The older access-key/secret-key pair is not enough, and the node tells you so rather than failing obscurely. Requests go to /text-to-video/kling-3.0-turbo on the unified /tasks API, which means Task Status can poll the same job if your workflow got interrupted.
Get the key from https://kling.ai/dev. New accounts need KYC activated first, which is the actual friction in this whole exercise - not the install.
Install
ComfyUI Manager → Install Custom Nodes → search "Kling Direct" → Install → restart. Or:
cd ComfyUI/custom_nodes
git clone https://github.com/IxMxAMAR/ComfyUI-Kling-Direct
The pack has no model downloads and no heavy dependencies - stdlib plus requests, Pillow, numpy, torch and opencv-python, all already present in a working ComfyUI. Your "install" is really account setup, and the pack ships a Kling API Health Check node that pings the free /account/costs endpoint so you can confirm auth before you spend anything. If you're outside the Singapore region, insert Kling Region Selector after Auth: china and us are options, and a wrong region fails in a way that looks like a key problem.
What to expect once you press queue
Text to video is the slowest thing in this pack because there's no reference image to anchor on - the model is generating the whole scene. The node blocks and polls, with intervals backing off to 15 seconds after two minutes and 30 after five, and a 1200-second ceiling before it raises a timeout. Cancelling mid-poll propagates within about a second, so you're not trapped; you just don't get the credits back.
Expect quality to be iterative, not first-try. Closed video models reward specific, physical prompts - subject, action, camera, light - and punish adjective piles the same way local models do. Start at 720p and 5 seconds, get a take you like, then spend the 1080p money. And check Cost Estimator before queueing a batch of storyboard variants, because the per-call pricing on this tier adds up fast enough that nobody wants to discover it from their card statement.
Inputs (5)
| Name | Type | Default | Description |
|---|---|---|---|
| auth | KLING_AUTH | — | |
| prompt | STRING | Text description of the video. Multi-shot format: 'shot 1, 3, words; shot 2, 2, words;' (shot number, seconds, prompt). | |
| resolution | COMBO | 720p | Output resolution. |
| aspect_ratio | COMBO | 16:9 | Output video aspect ratio. |
| duration | INT | 53–15 | Video duration in seconds. |
Outputs (5)
| Name | Type | Description |
|---|---|---|
| video | IMAGE | — |
| video_file | STRING | — |
| audio | AUDIO | — |
| url | STRING | — |
| task_id | STRING | — |