Create lip-synced talking videos using portrait images or videos
digitalhuman is a digital human video service for producing talking-head videos. It can create lip-synced speaking videos from facial images or videos. The creation interface accepts inputs such as audio, text, and voice identifiers, and delivers results as video URLs and task information. It is suited to production workflows with established character visuals and narration content, focusing on character expression rather than generating arbitrary scenes from scratch.
Input parameters and result formats vary by service. Use the public API for this model and follow its guide for generation, task retrieval and editing operations.
Specifications and API features
Creation method
Lip-synced video driven by facial images or facial videos
Video URL、task status、final video duration and width/height fields
The above are digitalhuman's creation capabilities and API fields. Final video dimensions and duration are subject to actual results.
Core capabilities
Make existing characters speak
Use a facial image or character video as the starting point and create lip-synced videos around narration content. It is suitable for projects with an established on-camera character that do not want to reshoot every time. Its core focus is talking characters; scripts, background design, and post-production packaging should still be planned separately in the production workflow.
Organize content around voice
The interface provides an audio URL as well as text and voice identifier fields, making it convenient to organize tasks around existing recordings or narration scripts. Before use, determine which narration material to use, and avoid treating filling in all fields at once as a general practice; validate specific input combinations with samples before using them at scale.
Support task-based production
Generation results include task identifiers, status, and video URLs, making them suitable for integration into content production workflows. Asynchronous options and callback addresses can be used to arrange result delivery; public MCP tools also provide task polling, batch retrieval, and deletion capabilities, allowing generation and subsequent organization to be handled separately.
Use Cases
Spoken Course Knowledge Points
Prepare presenter image assets and segmented scripts or recordings to create character narration clips for each knowledge point. The delivered videos can continue to be used for course editing, adding subtitles, and inserting presentation visuals. First verify the pronunciation and lip movements of specialized terms, then process the full set of content to better maintain consistency in the final video.
Product Explainer Videos
Organize product feature descriptions into spoken content and create introduction videos with a fixed presenter image. The output character narration clips can be combined with product screenshots, operation demonstrations, and brand assets. It handles digital human delivery; interface operations, subtitles, or full promotional video packaging should not be considered automatically included capabilities.
Fixed-Character Series Content
Continuously submit different narration content around the same character assets to create program openings, announcements, or knowledge-sharing clips. After generation, save the task identifier and video URL for archiving and later clip selection. Before proceeding in batches, first create a trial with representative content to check the voice, character performance, and editing compatibility.
How to Choose This Model
Choose It When You Have a Specific Presenter
If the task is to have a specified person image explain a piece of content, digitalhuman's image or video input and lip-sync method are better suited to the need. If the goal is complex camera movement, environmental changes, or multi-character storytelling, prioritize scene-generation video capabilities; do not treat character narration and general video generation as the same creative approach.
Decide on Assets First, Then the Workflow
When you already have a recording, you can organize the task around the audio URL; when starting from a script, you can try a combination of text and a voice identifier. If you need a customized voice, you can separately prepare it using MCP's voice cloning capability. Do not interpret the service as different performance tiers based on the old engine field; use the actual character assets and narration results to determine the production plan.
Get Started
Select an Image or Source Video
Provide image_url or video_url to define the character; for audio, use an existing audio_url, or use text with an already created voice_id, without confusing audio preparation with video generation.
Submit a Lip-Sync Task
Provide the character and narration input to /digital-human/videos, and keep the default control values initially; engine and resolution are deprecated, and new projects should integrate using the current default settings.
Review Segment by Segment Against the Audio
Save the task_id asynchronously, then query it through the task guide or receive callbacks; check mouth edges, pauses, and continuity, and create subtitles only after reading the final video's actual width, height, and duration.
Trial Recommendation: Revising Speech with Existing Footage
Input and Goal
Use a clear front-facing video of a single person and new narration audio, preserving the person and background while aligning the lip movements with the new speech.
Acceptance and Next Steps
First prepare the source video and audio_url; if using the text-based option, voice_id is also required. Check lip edges, pauses, and frame-to-frame continuity, and do not promote differences in the deprecated engine.
Usage Boundaries
Inputs center on face images or videos; this does not mean that any image can produce stable speech footage. It is recommended to first test the selected subject material, checking mouth performance, occluded areas, and visual continuity before using it for official content; multiple people in the same frame or complex movements should also not be handled directly through the single-person narration workflow.
The engine and resolution fields have been marked deprecated, so it is not recommended to rely on them when planning versions or image quality for new projects. The actual width, height, and duration of the finished video can be read from the response; when a fixed delivery size is required, post-production cropping, scaling, and acceptance should be included in the production workflow.
Voice cloning is a supporting MCP feature, and audio_url in the generation interface cannot be used directly as the entry point for cloning samples. Appropriate authorization should be obtained before using a person's likeness and voice, and the three stages of reference voice preparation, narration audio production, and digital human video generation should be distinguished.
Frequently Asked Questions
Is digitalhuman a general-purpose text-to-video model?
It is mainly used for lip-synced talking based on face images or videos. The text field is for narration content and should not be understood as being able to generate arbitrary videos solely from scene descriptions. To create complex environments, shots, or storylines, choose a more suitable creation method.
Must the character material be a video?
It does not have to be limited to video; digital human creation also supports starting from face images. The interface provides image_url and video_url respectively, which can be prepared according to the material type. For first-time use, it is recommended to make a short test clip first, confirm the character's performance, and then use the same material for subsequent tasks.
Can I use my own recording or script?
The interface provides audio_url, as well as text and voice_id fields, allowing creation based on recordings or scripts. The two approaches should not be mixed without validation; first choose the narration method, test the corresponding material combination, and then use the successful configuration for ongoing production.
Can I clone a voice directly in the generation request?
The accompanying MCP tool supports cloning a voice using a short reference audio clip, but the video generation interface does not have a dedicated field for cloning samples. Prepare the voice separately first, then arrange talking-head generation; do not confuse character-driving audio with voice-cloning reference material.
How do I obtain the video after generation?
Submit a task through POST /digital-human/videos; the response provides the video URL and task-related fields. When using an asynchronous workflow, you can use callbacks or MCP task queries to obtain progress, and determine from the returned status whether a usable result is available before downloading the video for post-production.