Skip to main content
Kling Lip Sync re-animates one detected face in a source video to match a speech track. AnyFast exposes the native Kling workflow: first identify the face, then submit session_id and face_choose at the root of the lip-sync request.
The lip-sync audio input is either face_choose[].audio_id or face_choose[].sound_file. To start from text, first call the Kling Digital Human TTS endpoint, then pass the returned audio id as face_choose[].audio_id.

Workflows

voice_id selects a TTS voice. audio_id identifies the audio generated by TTS. Do not pass a voice_id where the lip-sync endpoint expects an audio_id.

Prerequisite: identify the face

Call POST /kling/v1/videos/identify-face with either video_url or video_id. Save:
The source video must be 2–60 seconds, MP4 or MOV, 720p or 1080p, no more than 100 MB, and between 512 px and 2,160 px on each side. The complete request is documented in the Face Recognition tab of the Kling Lip Sync API Reference.

Create a lip-sync task

cURL
The endpoint currently supports one face_choose item. Provide exactly one of audio_id and sound_file.

Timing and volume

The cropped audio must be 2–60 seconds. For sound_file, use a public URL or raw Base64. MP3, WAV, and M4A have been validated through AnyFast; keep files within 5 MB.

Query the result

cURL
Poll data.task_status through submitted and processing. Stop on succeed or failed. A successful response returns the result at data.task_result.videos[0].url and its reported duration at data.task_result.videos[0].duration.

Kling Lip Sync API Reference

Review request, response, timing, and query fields.