Kling Lip Sync API (Kling Lip Sync)
Make an existing Kling video (5 seconds or 10 seconds) "speak" according to audio or text — that is, lip sync. Combined with the image2video feature of /kling/videos (to make a photo move), it can form a complete "talking photo / digital human presentation" workflow.
This API is a single-step convenient wrapper provided by qiyaov, designed for common audio/text-driven scenarios; it is not a field mirror of Kling's official multi-step "Face Recognition → Advanced Lip Sync" API. Please refer to the parameter table on this page.
- API Endpoint:
POST https://api.qiyaov.com/kling/lip-sync
- Request Format:
application/json
- Response Format:
application/json
- Billing: 2.45 Credits per successful call (fixed)
| Field |
Value |
Description |
authorization |
Bearer ${API_KEY} |
Your API key, get it here |
content-type |
application/json |
Request body format |
accept |
application/json |
Response format |
¶ Request Parameters (Request Body)
| Parameter |
Type |
Required |
Default |
Description |
mode |
string |
Yes |
— |
Generation mode. Enum: audio2video (audio-driven), text2video (text-driven) |
video_id |
string |
One of two |
— |
ID of a Kling-generated video (for example, the video_id returned by /kling/videos image2video). Only supports 5s/10s videos generated within 30 days. Choose one between video_id and video_url; they cannot be passed at the same time |
video_url |
string |
One of two |
— |
Publicly accessible video link. Constraints: .mp4/.mov, ≤100MB, duration 2–10s, 720p/1080p only, side length 720–1920px. Choose one between this and video_id |
audio_url |
string |
Conditional |
— |
Download URL of the driving audio, required when audio2video + audio_type=url. Formats .mp3/.wav/.m4a/.aac, ≤5MB |
audio_type |
string |
No |
url |
Audio transfer method. Enum: url, file (effective when using audio2video) |
audio_file |
string |
Conditional |
— |
Base64 of the audio file, required when audio_type=file. Same formats as above, ≤5MB |
text |
string |
Conditional |
— |
Text to be read aloud, required when using text2video, up to 120 characters |
voice_id |
string |
Conditional |
— |
Voice ID, required when using text2video |
voice_language |
string |
No |
zh |
Voice language. Enum: zh, en (effective when using text2video) |
voice_speed |
float |
No |
1.0 |
Speech speed, range 0.8–2.0, accurate to one decimal place (effective when using text2video) |
callback_url |
string |
No |
— |
Callback address. Passing this or async=true enables asynchronous mode: immediately returns task_id, then calls back after the result is generated |
async |
boolean |
No |
false |
Whether to run asynchronously. When true, immediately returns task_id, used with /kling/tasks polling or callback_url callback |
¶ Request Examples
¶ 1) Audio-driven (audio2video)
curl -X POST 'https://api.qiyaov.com/kling/lip-sync' \
-H 'authorization: Bearer ${API_KEY}' \
-H 'content-type: application/json' \
-d '{
"mode": "audio2video",
"video_id": "895055164389466178",
"audio_url": "https://cdn.acedata.cloud/6f7d62b18b.wav"
}'
¶ 2) Text-driven (text2video)
curl -X POST 'https://api.qiyaov.com/kling/lip-sync' \
-H 'authorization: Bearer ${API_KEY}' \
-H 'content-type: application/json' \
-d '{
"mode": "text2video",
"video_id": "895055164389466178",
"text": "哥,好久不见,我一切都好,你要照顾好自己。",
"voice_id": "genshin_vindi2",
"voice_language": "zh",
"voice_speed": 1.0
}'
¶ Response Example (Synchronous Success)
{
"success": true,
"task_id": "07a3ec65-9f7e-4a09-b7b7-282684082527",
"video_id": "895055968777281546",
"video_url": "https://cdn.acedata.cloud/assets/examples/kling/6c68c267-065b-4423-b66b-a0e4c59ee0d5-6a664a591a53.mp4",
"duration": "4.966",
"state": "succeed"
}
| Field |
Type |
Description |
success |
boolean |
Whether successful |
task_id |
string |
ID of this task (can be used for /kling/tasks queries) |
video_id |
string |
Kling ID of the generated video (can be used as input for the next extend/lip-sync) |
video_url |
string |
URL of the generated talking video (transferred to this platform's CDN, valid long-term) |
duration |
string |
Video duration (seconds) |
state |
string |
Task status: succeed / failed |
¶ Asynchronous Mode and Queries
When callback_url or async: true is passed, the API immediately returns task_id; afterwards you can:
- Poll:
POST /kling/tasks, body { "action": "retrieve", "id": "<task_id>" } (free)
- Callback: After generation is complete, the result is POSTed to your
callback_url
¶ Complete Workflow: Talking Photo (image2video → lip-sync)
# Step 1: Animate the photo and obtain video_id
curl -X POST 'https://api.qiyaov.com/kling/videos' \
-H 'authorization: Bearer ${API_KEY}' -H 'content-type: application/json' \
-d '{"model":"kling-v2-1-master","action":"image2video","start_image_url":"https://cdn.acedata.cloud/4hfydw.jpg","prompt":"look at camera, natural","duration":5,"mode":"pro"}'
# → { "video_id": "895055164389466178", ... }
# Step 2: Lip-sync with audio
curl -X POST 'https://api.qiyaov.com/kling/lip-sync' \
-H 'authorization: Bearer ${API_KEY}' -H 'content-type: application/json' \
-d '{"mode":"audio2video","video_id":"895055164389466178","audio_url":"https://cdn.acedata.cloud/assets/examples/fish/5ade0339-5f11-487e-aacc-06a908271706-8e3fcb0e5547.mp3"}'
# → { "video_url": "https://cdn.acedata.cloud/assets/examples/kling/6c68c267-065b-4423-b66b-a0e4c59ee0d5-6a664a591a53.mp4", ... }
¶ Error Response
{
"success": false,
"error": { "code": "bad_request", "message": "one of video_id or video_url is required" },
"trace_id": "f07cab09-3c18-4d74-9030-64ee840d9f16",
"task_id": "f490537f-2e5c-4739-8149-6252fba2091c"
}
| HTTP |
code |
Meaning |
| 400 |
bad_request |
Missing or invalid parameters (e.g., mode not provided, conflict between video and audio options, text exceeds 120 characters) |
| 401 |
authorization_missing |
Missing or invalid API key |
| 403 |
forbidden |
Content was blocked by risk control |
| 429 |
too_many_requests |
Upstream concurrency limit; please retry later |
| 500 |
api_error |
Upstream or internal error |
¶ Notes
video_id must be a Kling video generated within 30 days, and must be 5s or 10s; otherwise, use video_url to provide a video that meets the constraints.
- The input video is recommended to feature a clear frontal face and a single person for the best lip-sync results.
- The audio/text duration should match the video duration (audio must not exceed the video length).
- Billing occurs upon success (2.45 Credits/request); parameter validation failures (4xx) are not billed.