Short video model with first and last frame control and synchronized audio
kling-v2-6 is Kuaishou Kling's V2.6 video model, suitable for turning text concepts or static images into short videos. Its value lies in the pro mode, which provides both first and last frame control and synchronized audio, making it suitable for product showcases, atmospheric clips, and story shots; it can also be used in talking photo workflows to create lip-sync videos by combining portrait photos with existing audio.
Input parameters and result formats vary by service. Use the public API for this model and follow its guide for generation, task retrieval and editing operations.
Specifications and API features
Clarify capacity, inputs and outputs, and invocation methods before selecting a model.
Creation methods
Text-to-video, image-to-video; also includes a lip-sync workflow using photos and audio
Video duration
The platform supports 5 or 10 seconds
Generation modes
std, pro
Aspect ratios
16:9, 9:16, 1:1
First and last frame control
pro image-to-video supports a first frame plus a last frame; the last frame cannot be used independently
Synchronized audio
pro supports generate_audio; disabled by default
Result delivery
Video link, video ID, task ID, and status; supports asynchronous processing and callbacks
The above are this platform's V2.6 invocation specifications. Photo lip-sync is a separate workflow and is not the same as synchronized audio during video generation.
Core Capabilities
Learn what kling-v2-6 can bring to your work.
Generate Sound and Visuals Together
V2.6's pro mode supports enabling accompanying audio while generating videos, making it suitable for short-form creations that need both visuals and sound. Prompts can be organized around the subject, action, and scene, then audio can be enabled through generate_audio; when only visual assets are needed, keep it off for easier dubbing and editing later.
Structure Shots with Start and End Points
Image-to-video can begin from a specified first frame, and pro mode can also add a final frame to set clear opening and ending visuals for a clip. It is suitable for projects with existing product images, character images, or storyboards, allowing creation to develop around a selected composition; first and last frames constrain the visual boundaries rather than providing frame-by-frame editing of every intermediate frame.
Turn Portrait Photos into Talking Clips
The talking photo workflow accepts a portrait image and existing audio, first animates the photo, then lip-syncs it to the audio, delivering the final video in a single submission. Use the prompt to describe actions and expressions during the animation stage, and explicitly select kling-v2-6; this workflow is suited to tasks with existing recordings rather than synthesizing speech directly from text.
Use Cases
Start with specific tasks to find where the model can be effective.
Product Showcases and Social Short Videos
Enter a product hero image and a description of the display action to create short videos suitable for landscape, portrait, or square publishing. If the ending visual has already been designed, use pro to add a final frame so the product showcase lands on the intended composition; the delivered video link can enter the editing workflow, where brand text, subtitles, and a complete marketing arrangement can be added.
Sound-Enabled Atmosphere and Story Shots
Start with descriptions of the scene, character actions, and events, use text-to-video to create short shots, and enable synchronized accompanying audio in pro mode. This is suitable for first validating a story segment or sound-enabled atmosphere concept before deciding whether to move into formal editing; organizing the task as a clear single clip is easier to review than cramming an entire long-form script into one generation.
Talking Portraits with Existing Recordings
Prepare a clear front-facing photo of one person along with a recording that matches the target video length, then submit them through the talking photo entry point. The output includes the final lip-synced video, and you can also obtain the video link for the photo animation stage, making it easier to separately review character motion and lip-sync quality. It is suitable for character introductions, brief explanations, and talking-head samples.
How to Choose This Model
Choose based on task complexity, input materials, and expected results.
Choosing Between V2.5 Turbo and V2.6
If you need first and last frames, the pro modes of both can be considered; if you want to obtain audio at the same time as generation, V2.6 pro better meets the task requirements, as V2.5 Turbo does not provide this feature. For clips that do not require sound or a last frame, start with V2.6 std; choose pro when sound or a specific ending shot is needed, rather than judging solely by the version name.
When to Switch to V3 or Omni
V2.6 is suitable for fixed short durations, text- or first-frame-driven generation, and audio creation in pro mode. If the task requires integer durations of 3–15 seconds, native 4K mode, or dedicated camera movement controls, consider the corresponding V3 capabilities; if you need to combine multiple reference images, reference videos, or directly edit existing videos, choose O1 or V3 Omni rather than applying these capabilities to V2.6.
Get Started
From a small-scale task to production integration.
01
Prepare Tasks and Materials
Define the goal, required inputs, and output requirements, using real business examples as a starting point.
02
Try It in the API Debugging Area
Open the trial page, confirm the parameters supported by this entry point, then submit a small-scale task to review the results.
03
Integrate According to the API Documentation
Keep the complete model ID, use the request format specified in the documentation, and confirm billing rules on the Pricing page.
Usage Limits
Understand output quality and capability boundaries before formal use.
std does not support synchronized audio or last frames; choose pro to use these two features. Image-to-video requires a first-frame image, and a last frame can only be used together with a first frame; before submitting, clearly describe the desired action sequence to avoid mistaking start and end frame constraints for precise control over the entire motion path.
V2.6 does not support mode=4k or dedicated camera movement parameters through camera_control. You can express camera intentions in the prompt, but this differs from parameterized camera movement control; multi-image references, reference videos, and video editing are workflows for other models, not general input capabilities of this model.
Photo lip-sync requires accessible image and audio links, and clear front-facing images of a single person are recommended. Supported audio formats include mp3, wav, m4a, and aac; files must not exceed 5MB, and the duration is recommended not to exceed the target video; this workflow relies on existing audio and does not replace voice creation or long-form spoken content production.
Frequently Asked Questions
Answers to common questions about using kling-v2-6.
How should I choose between V2.6 std and pro?
If you are only creating clips without synchronized audio and do not need to specify an end frame, you can choose std. If you need synchronized audio or start-and-end frame control, use pro. Both modes support 5-second or 10-second videos; these pro features must be enabled explicitly and are not all automatically turned on simply by selecting the mode.
How do I enable synchronized audio in V2.6?
In the /kling/videos request, explicitly set model to kling-v2-6, mode to pro, and generate_audio to true. This switch is disabled by default. Synchronized audio and submitting pre-recorded audio for photo lip-sync are two different methods; choose based on whether you already have a recording.
Can I generate a video using only an end-frame image?
You cannot use an end frame alone as input for this image-to-video workflow. Use action=image2video, provide start_image_url, and add end_image_url as needed in pro mode. If you only have one image, use it as the start frame and describe the desired action in the prompt.
How do I make a photo speak along with a recording using V2.6?
Submit image_url and audio_url to /kling/talking-photo, explicitly specify model=kling-v2-6, and choose 5 seconds or 10 seconds. The prompt can describe the movements and expressions during the photo animation stage. The final video_url is the lip-synced video, while source_video_url is the intermediate photo animation video.
How do I receive generation results without keeping the connection waiting?
You can set async=true, save the returned task_id, and then retrieve the status and results through task queries; you can also configure callback_url to receive completion notifications. After obtaining video_url, download it or proceed to the editing workflow. Both generation endpoints should explicitly specify kling-v2-6 to avoid using other default models.
Model information · Updated: 2026-10-01. See the API and pricing sections for request parameters and billing rules.