Kling Videos Generation API Integration Guide

This article will introduce a Kling Videos Generation API integration guide, which can generate Kling official videos by inputting custom parameters.

Application Process

To use the Kling Videos Generation API, first go to the qiyaov Console to obtain your API Token and keep it for later use.

If you have not yet logged in or registered, you will be automatically redirected to the login page to register and log in. After completion, you will automatically return to the current page.

One API Token can call all services on the platform; there is no need to apply separately for each service. Your first application will receive free credits for a free trial; when credits are insufficient, you can top up your general balance in the Console.

📘 Full documentation: Kling Videos Generation API →

Basic Usage

First, let's understand the basic usage method. By inputting the prompt prompt, generation action action, first-frame reference image start_image_url, and model model, you can obtain the processed result. First, you need to simply pass an action field with the value text2video. It mainly includes three actions: text-to-video (text2video), image-to-video (image2video), and video extension (extend). Then, we also need to input the model model. Currently, there are mainly the kling-v1, kling-v1-6, kling-v2-master, kling-v2-1-master, kling-v2-5-turbo, kling-v2-6, kling-v3, kling-v3-omni, and kling-o1 models. The specific content is as follows:

Here, you can see that we have set the Request Headers, including:

  • accept: The format of the response result you want to receive. Enter application/json here, which is JSON format.
  • authorization: The key for calling the API. After applying, you can directly select it from the dropdown.

Additionally, the Request Body is set, including:

  • model: The model for generating videos. It mainly includes the kling-v1, kling-v1-6, kling-v2-master, kling-v2-1-master, kling-v2-5-turbo, kling-v2-6, kling-v3, kling-v3-omni, and kling-o1 models.
  • mode: The mode for generating videos. Optional values are standard mode std, fast mode pro, and native 4K mode 4k. Among them, 4k only supports kling-v3 and kling-v3-omni, and is incompatible with camera_control (camera movement control).
  • action: The action of this video generation task. It mainly includes three actions: text-to-video (text2video), image-to-video (image2video), and video extension (extend).
  • start_image_url: When selecting the image-to-video action image2video, the uploaded first-frame reference image link is required.
  • end_image_url: Optional for image-to-video; specifies the final frame.
  • duration: Video duration, in seconds. kling-v3 and kling-v3-omni support integer durations of 3-15 seconds; kling-o1 only supports 5 seconds; other models support 5 or 10 seconds.
  • generate_audio: Whether to generate audio simultaneously. Optional, Boolean value. Supported by kling-v3, kling-v3-omni, and kling-v2-6 (pro mode only). Defaults to false.
  • aspect_ratio: Video aspect ratio. Optional, supports 16:9, 9:16, and 1:1, with a default of 16:9.
  • cfg_scale: Relevance strength, ranging from [0,1]. The larger the value, the more closely it matches the prompt.
  • camera_control: Optional object parameters for controlling camera movement, supporting type/simple presets and configurations such as horizontal, vertical, pan, tilt, roll, and zoom.
  • negative_prompt: Optional negative prompt for content you do not want to appear, up to 200 characters.
  • image_list: Omni reference image list, applicable to the kling-o1 and kling-v3-omni models. See "Omni Universal Reference" below for usage.
  • video_list: Omni reference video list (supports video editing), applicable to the kling-o1 and kling-v3-omni models. See "Omni Universal Reference" below for usage.
  • prompt: Prompt.
  • callback_url: The URL that needs to receive callback results.
  • async: Optional. When set to true, the API immediately returns a task_id, without needing to provide callback_url; subsequently, obtain results by polling through the corresponding task query API.

After selecting, you can find that the corresponding code is also generated on the right, as shown in the image:

Click the "Try" button to test, as shown above. Here, we obtain the following result:

{
  "success": true,
  "video_id": "900798310464749610",
  "video_url": "https://cdn.acedata.cloud/assets/examples/kling/6c68c267-065b-4423-b66b-a0e4c59ee0d5-6a664a591a53.mp4",
  "duration": "5.041",
  "state": "succeed",
  "task_id": "6c68c267-065b-4423-b66b-a0e4c59ee0d5"
}

The returned result contains multiple fields, described as follows:

  • success: The status of the video generation task at this time.
  • task_id: The ID of the video generation task at this time.
  • video_id: The video ID of the video generation task at this time.
  • video_url: The video link of the video generation task at this time.
  • duration: The video link duration of the video generation task at this time.
  • state: The status of the video generation task at this time.

You can see that we have obtained satisfactory video information. We only need to retrieve the generated Kling video according to the video link address in data in the result.

Additionally, if you want to generate the corresponding integration code, you can copy it directly. For example, the CURL code is as follows:

curl -X POST 'https://api.qiyaov.com/kling/videos' \
-H 'accept: application/json' \
-H 'authorization: Bearer {token}' \
-H 'content-type: application/json' \
-d '{
  "action": "text2video",
  "model": "kling-v3",
  "prompt": "White ceramic coffee mug on glossy marble countertop with morning window light. Camera slowly rotates 360 degrees around the mug, pausing briefly at the handle."
}'

Model Capability Matrix

Different models vary significantly in their support for parameters. The following matrix is compiled from the official Kling video models documentation. Before calling, please verify whether the current model / mode / duration combination supports the features you need; otherwise, errors such as model/mode/duration(...) is not supported with image_tail will be returned.

Model Mode end_image_url (First/Last Frame) generate_audio (Audio) camera_control (Camera Movement) Notes
kling-v1 std / pro ✅ Only duration=5 ❌ ✅ Only duration=5 extend does not support negative_prompt or cfg_scale
kling-v1-6 std ❌ ❌ ❌ Multi-image-to-video, all extend modes available
kling-v1-6 pro ✅ ❌ ❌
kling-v2-master — ❌ ❌ ❌ Single mode, only duration=5/10
kling-v2-1-master — ❌ ❌ ❌ Single mode, only duration=5/10
kling-v2-5-turbo std ❌ ❌ ❌
kling-v2-5-turbo pro ✅ ❌ ❌
kling-v2-6 std ❌ ❌ ❌
kling-v2-6 pro ✅ ✅ ❌ The only non-v3 model that supports audio simultaneously
kling-v3 std / pro ✅ ✅ ✅ duration range: 3–15 seconds
kling-v3 4k ✅ ✅ ❌ 4K mode is incompatible with camera movement
kling-v3-omni std / pro / 4k ✅ ✅ ❌
kling-o1 std / pro ✅ ❌ ❌ Only supports duration=5

Notes:

  • mode=4k is supported only by kling-v3 and kling-v3-omni; it is mutually exclusive with camera_control (camera movement).
  • end_image_url can only be used together with start_image_url when action=image2video. Passing only end_image_url (without start_image_url) will be rejected.
  • kling-v3 / kling-v3-omni accept any integer duration from 3–15 seconds; kling-o1 only accepts 5; all other models only accept 5 or 10.
  • generate_audio defaults to false. It is supported only by kling-v3, kling-v3-omni, and kling-v2-6 (pro mode).

Video Extension Feature

If you want to continue generating an already generated Kling video, you can set the parameter action to extend, and enter the ID of the video that needs to be continued. The video ID can be obtained according to the basic usage, as shown in the image below:

At this point, you can see that the video ID is:

"video_id": "030bb06d-98d4-4044-9042-0aa0822e8c8c"

Note that the video_id in the video here is the ID of the generated video. If you do not know how to generate a video, you can refer to the basic usage above to generate one.

Next, we must fill in the prompt for the next extension step to customize the generated video, and can specify the following content:

  • model: The model for generating videos, mainly including the kling-v1, kling-v1-5, and kling-v1-6 models.
  • mode: The mode for generating videos. Optional values are standard mode std, fast mode pro, and native 4K mode 4k (supported only by kling-v3 and kling-v3-omni, incompatible with camera movement control).
  • duration: The video duration for this video generation task, mainly including 5s and 10s.
  • start_image_url: When selecting the image-to-video behavior image2video, the uploaded first-frame reference image link is required.
  • prompt: Prompt.

The filling example is as follows:

After filling it in, the code is automatically generated as follows:

The corresponding Python code:

import requests

url = "https://api.qiyaov.com/kling/videos"

headers = {
    "accept": "application/json",
    "authorization": "Bearer {token}",
    "content-type": "application/json"
}

payload = {
    "action": "extend",
    "model": "kling-v1",
    "video_id": "030bb06d-98d4-4044-9042-0aa0822e8c8c",
    "prompt": "White ceramic coffee mug on glossy marble countertop with morning window light. Camera slowly rotates 360 degrees around the mug, pausing briefly at the handle.",
    "duration": 10
}

response = requests.post(url, json=payload, headers=headers)
print(response.text)

Click Run, and you can find that you will get a result as follows:

{
  "success": true,
  "video_id": "bbc3b105-ac72-4de2-8390-0cb37dc7d41e",
  "video_url": "https://cdn.acedata.cloud/assets/examples/gemini/04a043bd-6b23-4b4e-945c-ce48158c3eee-3a89912507c7.mp4",
  "duration": "9.6",
  "state": "succeed",
  "task_id": "3ece87e6-3ee3-4f5e-bd70-5ae5eca89a23"
}

As can be seen, the result content is consistent with the above, which implements the video extension feature.

Omni Universal Reference (Video Editing / Reference Video / Multi-Image Reference)

kling-o1 and kling-v3-omni are two independent models, both supporting the "universal reference" capability. Based on text-to-video (action=text2video), you can additionally pass in reference images or reference videos to achieve multi-image reference, reference video, and directly editing existing videos.

Core convention: Reference materials must be referenced in prompt in the form of <<<image_1>>>, <<<video_1>>> (numbering starts from 1), referring to the materials at the corresponding positions in image_list / video_list, before the model will apply these references. If materials are only passed but not referenced in the prompt, the materials will be ignored.

Security note: The current API does not expose element_list. The IDs of the Kling Element Library are not tenant-isolated. Before providing a tenant-isolated Element Management API, please use image_list to pass in subject reference images.

Omni requests do not support negative_prompt, cfg_scale, or camera_control, and cannot use mode=4k. When reference videos are included, generate_audio must be false.

Reference Video and Video Editing (video_list)

video_list is used to pass reference videos and is the most commonly used scenario for this capability. The array element fields are as follows:

  • video_url: Reference video link, cannot be empty. Up to 1 MP4/MOV video, file size ≤200MB, frame rate 24–60fps. kling-o1 requires a duration of 3–10 seconds and both width and height of 700–2160px; kling-v3-omni requires a duration of 3–15.5 seconds, both width and height of 700–4553px, total pixels ≤8,294,400, and an aspect ratio of 0.4–2.
  • refer_type: Reference type, optional base (default, the base video to be edited, meaning "edit the video directly," allowing elements to be added, removed, or modified, composition to be changed, style to be changed, colors to be changed, weather to be changed, etc.) or feature (feature reference, referencing its style / camera movement / continuation into the next shot).
  • keep_original_sound: Whether to preserve the original video audio, optional yes (preserve) or no (remove).

Note: When a reference video exists, generate_audio must be false. A video with refer_type=base cannot additionally specify a first frame / last frame.

A CURL example for editing an existing video (changing the video to an anime style) is as follows:

curl -X POST 'https://api.qiyaov.com/kling/videos' \
-H 'accept: application/json' \
-H 'authorization: Bearer {token}' \
-H 'content-type: application/json' \
-d '{
  "action": "text2video",
  "model": "kling-o1",
  "mode": "std",
  "duration": 5,
  "prompt": "Change <<<video_1>>> to a cinematic anime style, preserving the original motion and composition",
  "video_list": [
    {
      "video_url": "https://cdn.acedata.cloud/your-reference-video.mp4",
      "refer_type": "base",
      "keep_original_sound": "no"
    }
  ]
}'

Multi-image reference (image_list)

image_list is used to pass reference images (elements / scenes / styles, etc.). The array element fields are as follows:

  • image_url: Reference image link, cannot be empty. Requirements: .jpg/.jpeg/.png format; file size ≤10MB; shortest side ≥300px; aspect ratio 1:2.5 ~ 2.5:1.
  • type: Optional. When not passed, it is used as a pure reference image; when first_frame / end_frame is passed, it is used as the first frame / last frame respectively (equivalent to start_image_url / end_image_url).

When using it, reference it in prompt with <<<image_1>>>, <<<image_2>>>. Quantity limits: when no reference video exists, reference images ≤ 7; when a reference video exists, reference images ≤ 4. When only passing the first / last frame, start_image_url / end_image_url can also be used directly, but the last frame must be used together with the first frame.

Note: If start_image_url / end_image_url and image_list are passed at the same time, the first / last frames will be placed before image_list, which may affect the index correspondence of <<<image_N>>>. It is recommended to choose one: when first / last frames are needed, specify them using type directly in image_list, and do not mix them with start_image_url / end_image_url.

A CURL example for generating a video with multi-image references:

curl -X POST 'https://api.qiyaov.com/kling/videos' \
-H 'accept: application/json' \
-H 'authorization: Bearer {token}' \
-H 'content-type: application/json' \
-d '{
  "action": "text2video",
  "model": "kling-o1",
  "mode": "std",
  "duration": 5,
  "prompt": "Make the character in <<<image_1>>> stand in the scene of <<<image_2>>>, with cinematic lighting",
  "image_list": [
    { "image_url": "https://cdn.acedata.cloud/subject.png" },
    { "image_url": "https://cdn.acedata.cloud/scene.png" }
  ]
}'

Asynchronous callback

Since the Kling Videos Generation API takes a relatively long time to generate, approximately 1–2 minutes, if the API does not respond for a long time, the HTTP request will remain connected, resulting in additional system resource consumption. Therefore, this API also provides support for asynchronous callbacks.

The overall process is: when the client initiates a request, it additionally specifies a callback_url field. After the client initiates the API request, the API will immediately return a result containing a task_id field, representing the current task ID. After the task is completed, the result of the generated video will be sent in POST JSON format to the callback_url specified by the client, which also includes the task_id field, so that task results can be associated through the ID.

Below, we will use an example to understand the specific operation.

First, a Webhook callback is a service that can receive HTTP requests. Developers should replace it with the URL of their own deployed HTTP server. For convenience of demonstration, a public Webhook example website https://webhook.site/ is used here. Opening this website will provide a Webhook URL, as shown in the figure:

Copy this URL, and it can be used as a Webhook. The example here is https://webhook.site/624b2c78-6dbd-4618-9d2b-b32eade6d8c3.

Next, we can set the callback_url field to the above Webhook URL and fill in the corresponding parameters. The specific content is shown in the figure:

Click Run, and you can see that a result is returned immediately, as follows:

{
  "task_id": "20068983-0cc9-4c6a-aeb6-9c6a3c668be0"
}

After waiting a moment, we can observe the generated video result at https://webhook.site/624b2c78-6dbd-4618-9d2b-b32eade6d8c3, as shown in the figure:

The content is as follows:

{
    "success": true,
    "video_id": "030bb06d-98d4-4044-9042-0aa0822e8c8c",
    "video_url": "https://cdn.acedata.cloud/assets/examples/gemini/04a043bd-6b23-4b4e-945c-ce48158c3eee-3a89912507c7.mp4",
    "duration": "5.1",
    "state": "succeed",
    "task_id": "20068983-0cc9-4c6a-aeb6-9c6a3c668be0"
}

You can see that there is a task_id field in the result. The other fields are similar to those above, and task association can be achieved through this field.

Error handling

When calling the API, if an error is encountered, the API will return the corresponding error code and information. For example:

  • 400 token_mismatched: Bad request, possibly due to missing or invalid parameters.
  • 400 api_not_implemented: Bad request, possibly due to missing or invalid parameters.
  • 401 invalid_token: Unauthorized, invalid or missing authorization token.
  • 429 too_many_requests: Too many requests, you have exceeded the rate limit.
  • 500 api_error: Internal server error, something went wrong on the server.

Error Response Example

{
  "success": false,
  "error": {
    "code": "api_error",
    "message": "fetch failed"
  },
  "trace_id": "2cf86e86-22a4-46e1-ac2f-032c0f2a4e89"
}

Conclusion

Through this document, you have learned how to use the Kling Videos Generation API to generate videos by entering prompt words and a first-frame reference image. We hope this document can help you better integrate and use this API. If you have any questions, please feel free to contact our technical support team.