All models

kling-v2-master

KuaishouVideo
Get your API key
kling-v2-master

Kling 2nd Generation Video Model for Short-Shot Creation

kling-v2-master is Kuaishou Kling's V2 Master video model, designed for short-shot creation starting from text concepts or a first-frame image, balancing visual quality and user experience. It is suitable for product showcases, concept storyboards, and animated photo assets, and can also handle animation tasks in talking photo workflows. When choosing, focus on single-shot expression rather than multi-asset editing or native audiovisual generation.

KuaishouModel Brand
VideoModel Type
Text · Image GuidedCreation Method
STANDARD APIs · QUICK SETUP

Bring this model into your workflow

Submit requests to the public API at api.acedata.cloud using the documented parameters, then use the results in your application.

API hostapi.acedata.cloud
modelkling-v2-master

Input parameters and result formats vary by service. Use the public API for this model and follow its guide for generation, task retrieval and editing operations.

Specifications and API Features

Creation Method
Text-to-video, first-frame image-to-video
Video Duration
Video generation on this platform: 5 or 10 seconds
Aspect Ratio Options
16:9, 9:16, 1:1
Model Control Scope
Single mode; does not support end frames or camera_control
Audio Workflow
Video generation does not support synchronized audio; talking photos use external audio
Photo Voiceover Input
Photo URL and audio URL; audio supports mp3/wav/m4a/aac, ≤5MB
Results and Tasks
Returns video link, task ID, and status; supports asynchronous processing and callbacks

The durations, aspect ratios, and asset requirements above correspond to this platform's supported usage scope. Talking photos are a combined creation workflow and are not equivalent to the model's native audio accompaniment.

Core Capabilities

From Text Concepts to Short Shots

Text-to-video is suitable for turning scene descriptions into watchable shot drafts. When creating, prompts can be organized around the subject, action, environment, and lighting, focusing one piece of content on a clear event. The balanced positioning of V2 Master is suitable for comparing different creative expressions, then using selected results for storyboard discussions or subsequent editing.

Use the First Frame to Establish a Visual Starting Point

Image-to-video uses an existing image as the opening foundation, then describes the desired action through prompts. Compared with starting entirely from text, this approach allows product images, portraits, or scene artwork to directly participate in creation. It emphasizes developing the visuals from a specified starting point and does not provide end-frame locking, making it better suited to short shots whose endings can develop freely.

Photo Animation and Voiceover Integration

The talking photo workflow accepts a portrait photo and existing audio, first animates the photo, then applies lip-sync processing according to the audio. After selecting V2 Master, it handles the animation generation portion. This workflow is suitable for character introductions or brief spoken segments with existing voiceovers, without mistakenly writing dialogue as an instruction for the video model to automatically generate sound.

Use Cases

Turn Static Product Images into Dynamic Assets

Input a product display image and describe a slight rotation, background atmosphere, or presentation action to generate dynamic clips for short-video editing. Choose landscape, portrait, or square aspect ratios according to the placement, then add selling-point captions and music during editing. It is recommended to highlight one presentation focus at a time rather than arranging an entire advertising storyline within the same shot.

Concept Storyboards and Shot Previsualization

Organize a single scene from the script into descriptions of the subject, action, and environment, then use text-to-video to create shot drafts; when art direction already exists, you can also start from a scene image. The deliverables are short clips that facilitate team discussion, suitable for comparing mood, composition, and action expression before deciding which content enters formal production or manual refinement.

Short Character Videos with Existing Voiceovers

Prepare a clear front-facing photo of one person and a voiceover, then use the talking photo entry point to create character introductions, character greetings, or brief explanations. The audio should stay within the selected video duration, and prompts should mainly describe expressions and actions. After completion, obtain the final video link and integrate it into existing publishing or editing workflows.

How to Choose This Model

How to Choose Between It and V2.1 Master

V2 Master is suitable for short-shot tasks driven by text or a first frame; V2.1 Master places greater emphasis on quality and consistency. Both use a single mode here, with durations of 5 or 10 seconds, and neither provides an end frame, audio, or structured camera movement. If visual consistency matters, compare results using the same material rather than treating version numbers as a direct guarantee of performance.

When to Choose Other Kling Versions

If the ending frame must be locked, consider V2.5 Turbo pro, which supports end frames; if synchronized audio is needed, choose V2.6 pro or V3. Multiple images, reference videos, and editing existing videos are better suited to O1 or V3 Omni. The value of choosing V2 Master lies in its clear text-to-video and first-frame creation workflow, not in covering these different control requirements.

Getting Started

Give a Single Shot a Clear Goal

Use text2video for text only; use image2video and start_image_url when you already have key visuals. Choose one subject action, and describe the lighting, composition, and appearance that should be preserved.

Use a Single Mode and Short Duration

Specify /kling/videos with model=kling-v2-master, use this model's single mode, and select a duration of 5 or 10 seconds. Image-to-video starts from the first frame, without adding an end frame or a multi-material list.

Review the Subject and Visual Quality

Use async or callback_url to retrieve task results, and save the task_id and video ID; focus on checking materials, subject deformation, and inter-frame continuity. The output does not include native audio, so voice-over and music should be produced separately.

Trial Recommendation: Quality-First Single-First-Frame Animation

Input and Goal

Use an image of a white ceramic cup as the first frame. The camera slowly pushes in, steam rises from the cup opening, and the background lighting remains stable.

Review and Next Steps

Review based on a short shot of 5 or 10 seconds. Do not provide an end frame or camera_control; if sound is needed, add voice-over separately or use a supported photo talking-head workflow.

Usage Limitations

  • V2 Master does not support end_image_url end-frame constraints or camera_control camera movement parameters. Prompts can express the desired action or camera feel, but this does not constitute structured control and cannot guarantee that the final image stops at a specified composition.
  • Video generation does not support generate_audio synchronized audio, nor does it support 4K mode. Talking photos require audio to be provided separately, and their sound comes from the input material; if dialogue, ambient sound, and visuals need to be generated together, choose a model with the corresponding capabilities.
  • Do not treat image_list or video_list as multi-asset creation capabilities of this model. Reference video editing and Omni multi-image reference belong to other models; V2 Master's image-to-video starts from the first-frame image, and complex asset combinations should be created in separate steps or with another model.

Frequently Asked Questions

Which model name should I use when calling it?

Explicitly specify model=kling-v2-master in /kling/videos or /kling/talking-photo. This clearly selects V2 Master rather than relying on the endpoint default. kling-v2-1-master is another model and should not be used as an alternative name for the same model.

What core inputs are needed for image-to-video?

Use /kling/videos, set action to image2video, provide start_image_url, and use prompt to describe the action and scene changes. Choose a duration of 5 or 10 seconds. This model does not support end-frame control, so the first frame should be treated as the starting point rather than constraining both the start and end points.

Can a photo directly speak the dialogue in the prompt?

You cannot use a video prompt as speech generation input. To create a talking photo, provide image_url and audio_url to /kling/talking-photo; prompt is used for actions or expressions during the photo animation stage. It is recommended to prepare the voiceover first, then have the photo perform lip-sync processing based on that audio.

Can it generate videos of any length?

This model generates videos with durations of 5 or 10 seconds, and talking photos also offer these two duration options. Longer content can be split into shots and then edited together; if you need more flexible single-generation durations, consider V3 instead of passing arbitrary seconds to V2 Master.

How do I integrate generation tasks into an application?

You can use async=true to obtain task_id and then query the task result; you can also set callback_url to receive completion notifications. The application should associate the task ID with the creation record and read state and video_url upon completion. Successful task submission and completed video generation are different stages and should not be treated as the same status.