Kling Lip Sync API (Kling Lip Sync)

Make an existing Kling video (5 seconds or 10 seconds) "speak" according to audio or text — that is, lip sync. Combined with the image2video feature of /kling/videos (to make a photo move), it can form a complete "talking photo / digital human presentation" workflow.

This API is a single-step convenient wrapper provided by qiyaov, designed for common audio/text-driven scenarios; it is not a field mirror of Kling's official multi-step "Face Recognition → Advanced Lip Sync" API. Please refer to the parameter table on this page.

  • API Endpoint: POST https://api.qiyaov.com/kling/lip-sync
  • Request Format: application/json
  • Response Format: application/json
  • Billing: 2.45 Credits per successful call (fixed)

Request Headers

Field Value Description
authorization Bearer ${API_KEY} Your API key, get it here
content-type application/json Request body format
accept application/json Response format

Request Parameters (Request Body)

Parameter Type Required Default Description
mode string Yes — Generation mode. Enum: audio2video (audio-driven), text2video (text-driven)
video_id string One of two — ID of a Kling-generated video (for example, the video_id returned by /kling/videos image2video). Only supports 5s/10s videos generated within 30 days. Choose one between video_id and video_url; they cannot be passed at the same time
video_url string One of two — Publicly accessible video link. Constraints: .mp4/.mov, ≤100MB, duration 2–10s, 720p/1080p only, side length 720–1920px. Choose one between this and video_id
audio_url string Conditional — Download URL of the driving audio, required when audio2video + audio_type=url. Formats .mp3/.wav/.m4a/.aac, ≤5MB
audio_type string No url Audio transfer method. Enum: url, file (effective when using audio2video)
audio_file string Conditional — Base64 of the audio file, required when audio_type=file. Same formats as above, ≤5MB
text string Conditional — Text to be read aloud, required when using text2video, up to 120 characters
voice_id string Conditional — Voice ID, required when using text2video
voice_language string No zh Voice language. Enum: zh, en (effective when using text2video)
voice_speed float No 1.0 Speech speed, range 0.8–2.0, accurate to one decimal place (effective when using text2video)
callback_url string No — Callback address. Passing this or async=true enables asynchronous mode: immediately returns task_id, then calls back after the result is generated
async boolean No false Whether to run asynchronously. When true, immediately returns task_id, used with /kling/tasks polling or callback_url callback

Request Examples

1) Audio-driven (audio2video)

curl -X POST 'https://api.qiyaov.com/kling/lip-sync' \
  -H 'authorization: Bearer ${API_KEY}' \
  -H 'content-type: application/json' \
  -d '{
    "mode": "audio2video",
    "video_id": "895055164389466178",
    "audio_url": "https://cdn.acedata.cloud/6f7d62b18b.wav"
  }'

2) Text-driven (text2video)

curl -X POST 'https://api.qiyaov.com/kling/lip-sync' \
  -H 'authorization: Bearer ${API_KEY}' \
  -H 'content-type: application/json' \
  -d '{
    "mode": "text2video",
    "video_id": "895055164389466178",
    "text": "哥,好久不见,我一切都好,你要照顾好自己。",
    "voice_id": "genshin_vindi2",
    "voice_language": "zh",
    "voice_speed": 1.0
  }'

Response Example (Synchronous Success)

{
  "success": true,
  "task_id": "07a3ec65-9f7e-4a09-b7b7-282684082527",
  "video_id": "895055968777281546",
  "video_url": "https://cdn.acedata.cloud/assets/examples/kling/6c68c267-065b-4423-b66b-a0e4c59ee0d5-6a664a591a53.mp4",
  "duration": "4.966",
  "state": "succeed"
}
Field Type Description
success boolean Whether successful
task_id string ID of this task (can be used for /kling/tasks queries)
video_id string Kling ID of the generated video (can be used as input for the next extend/lip-sync)
video_url string URL of the generated talking video (transferred to this platform's CDN, valid long-term)
duration string Video duration (seconds)
state string Task status: succeed / failed

Asynchronous Mode and Queries

When callback_url or async: true is passed, the API immediately returns task_id; afterwards you can:

  • Poll: POST /kling/tasks, body { "action": "retrieve", "id": "<task_id>" } (free)
  • Callback: After generation is complete, the result is POSTed to your callback_url

Complete Workflow: Talking Photo (image2video → lip-sync)

# Step 1: Animate the photo and obtain video_id
curl -X POST 'https://api.qiyaov.com/kling/videos' \
  -H 'authorization: Bearer ${API_KEY}' -H 'content-type: application/json' \
  -d '{"model":"kling-v2-1-master","action":"image2video","start_image_url":"https://cdn.acedata.cloud/4hfydw.jpg","prompt":"look at camera, natural","duration":5,"mode":"pro"}'
# → { "video_id": "895055164389466178", ... }

# Step 2: Lip-sync with audio
curl -X POST 'https://api.qiyaov.com/kling/lip-sync' \
  -H 'authorization: Bearer ${API_KEY}' -H 'content-type: application/json' \
  -d '{"mode":"audio2video","video_id":"895055164389466178","audio_url":"https://cdn.acedata.cloud/assets/examples/fish/5ade0339-5f11-487e-aacc-06a908271706-8e3fcb0e5547.mp3"}'
# → { "video_url": "https://cdn.acedata.cloud/assets/examples/kling/6c68c267-065b-4423-b66b-a0e4c59ee0d5-6a664a591a53.mp4", ... }

Error Response

{
  "success": false,
  "error": { "code": "bad_request", "message": "one of video_id or video_url is required" },
  "trace_id": "f07cab09-3c18-4d74-9030-64ee840d9f16",
  "task_id": "f490537f-2e5c-4739-8149-6252fba2091c"
}
HTTP code Meaning
400 bad_request Missing or invalid parameters (e.g., mode not provided, conflict between video and audio options, text exceeds 120 characters)
401 authorization_missing Missing or invalid API key
403 forbidden Content was blocked by risk control
429 too_many_requests Upstream concurrency limit; please retry later
500 api_error Upstream or internal error

Notes

  • video_id must be a Kling video generated within 30 days, and must be 5s or 10s; otherwise, use video_url to provide a video that meets the constraints.
  • The input video is recommended to feature a clear frontal face and a single person for the best lip-sync results.
  • The audio/text duration should match the video duration (audio must not exceed the video length).
  • Billing occurs upon success (2.45 Credits/request); parameter validation failures (4xx) are not billed.