Lip Sync
Make a portrait speak an audio track, or dub an existing video by re-syncing its lips to new speech, with Kling Avatar, OmniHuman, Seedance and more.
The Lip Sync node makes a face speak an audio track. Give it a portrait and a voice track, and it makes a talking-head video. Give it a video and new audio, and it re-syncs the lips in the video to the new speech, which dubs the clip. You choose the model, and the node shows the inputs that model needs.
When to use it
- A talking head from a single photo, for example an AI-generated character with a voiceover.
- A spokesperson video for a product demo.
- A dubbed version of a video: translated speech, re-synced to the speaker's lips.
- A performance directed by a prompt, such as singing into a microphone, with OmniHuman 1.5.
For a full avatar video from a script, with voices and captions built in, use AI Avatar.
Quick start
Add the node
Press Tab on the canvas and choose Video › Animate & Perform › Lip Sync.
Choose the model
Open the settings panel and choose a model under Provider. The node then shows the inputs that model needs: Portrait, Source video, or both.
Connect the face and the voice
Connect an image to Portrait, or a video to Source video, and an audio node to Audio, for example Text to Speech.
Run it
Click Run. The lip-synced video appears on the node.
Portrait or video
| What you connect | What the node makes | Models |
|---|---|---|
| A Portrait and Audio | A new talking-head video from a still image. | Kling Avatar, Kling Avatar Pro, InfiniTalk, OmniHuman 1.5, Seedance 2, Seedance 2 Fast, Seedance 2 Mini, Seedance 2.5, SadTalker |
| A Source video and Audio | The same video with its lips re-synced to the new audio. | Volcengine Lip Sync, Sync Lipsync v3, Lipsync 2 Pro, HeyGen Lipsync Precision, LatentSync, Video-Retalking |
| Either one, and Audio | Wav2Lip accepts a portrait or a video. | Wav2Lip |
Inputs
| Input | Accepts | What it does |
|---|---|---|
| Audio | Audio nodes | The speech or voiceover. Required for every model. |
| Portrait | Image nodes | A clear face photo, for portrait models. |
| Source video | Video nodes | The clip to dub, for dubbing models. |
The output, Video, is the URL of the lip-synced video. When several nodes are connected, the node card lets you choose which image, video and audio to use.
Settings
| Setting | What it does |
|---|---|
| Provider | The lip-sync model. The default is Kling Avatar. |
| Resolution | On Kling Avatar, Kling Avatar Pro, InfiniTalk and the Seedance models: 480p or 720p, with 720p as the default. Seedance 2 also offers 1080p. OmniHuman 1.5 offers 720p or 1080p, with 1080p as the default. |
| Motion Prompt (optional) | On the same models: head and expression motion, for example "slight head nods" or "expressive eyebrows". On OmniHuman 1.5 it directs the whole performance. |
| Pre & post text | Text always added before and after the motion prompt. See Prompt pre and post text. |
Settings for specific models
| Model | Settings |
|---|---|
| OmniHuman 1.5 | Motion Prompt directs the performance, for example "sing confidently into a microphone". Fast Mode trades some quality for speed. Seed makes a result repeatable; -1 means random. It animates people, pets and anime at any aspect ratio. |
| HeyGen Lipsync Precision | Dynamic Duration (on) adjusts the output length to the new audio. Remove Music Track (off) strips background music from the source video. Speech Enhancement (off) improves speech clarity. |
| Lipsync 2 Pro | Sync Mode sets what happens when the audio and the video have different lengths: Loop (the default), Bounce, Cut off, Silence or Remap. Temperature, from 0 to 1, sets how expressive the lip sync is; the default is 0.5. Active Speaker Detection syncs whoever is speaking. |
| Sync Lipsync v3 | Sync Mode, with Cut off as the default, and Active Speaker Detection, which you should turn on for a video with more than one person. The model manages its own expressiveness. |
| Volcengine Lip Sync | Mode: Lite for one frontal speaker, faster, or Basic for complex scenes with several speakers. Separate vocals (denoise) strips noise from the audio. Scene detection + speaker ID (Basic only) finds scene cuts and who speaks. Loop video if audio is longer (Lite only, on by default) and Reverse loop (ping-pong) handle audio that runs past the video. Template start time (seconds) sets where in the video the lips start to follow. The output is as long as the audio: a longer video is cut, and a shorter one loops. |
The Seedance models do native, phoneme-level lip sync in more than 8 languages. They work like Generate Video with the audio sent as reference audio.
Models
| Model | Maker | Modes | Credits | Details |
|---|---|---|---|---|
| minimax-h3 | MiniMax | Image to video, Text to video | from 230 | MiniMax Hailuo 3 — premium multimodal tier: first/last frame + image/video/audio references, native audio, 2K (default) or 768P output, 4-15s per-second pricing. |
| Kling Avatar Standard | Kuaishou | Lip sync | 280 | Lip-sync a still portrait to driving audio. Standard quality. |
| Kling Avatar Pro | Kuaishou | Lip sync | 560 | Premium lip-sync — better mouth shape and timing. |
| Seedance 2 | Bytedance | Image to video, Text to video | from 230 | Seedance 2 — premium tier with native audio. Per-second pricing by resolution. |
| Seedance 2 Fast | Bytedance | Image to video, Text to video | from 180 | Cheaper / quicker Seedance 2 tier. |
| Seedance 2 Mini | Bytedance | Image to video, Text to video | from 120 | Budget Seedance 2 tier — 480p/720p only, per-second pricing by resolution. |
| Seedance 2.5 | Bytedance | Image to video, Text to video | from 340 | Seedance 2.5 — up to 30s in one shot, native audio, wide multimodal references. 480p/720p/1080p. |
| OmniHuman 1.5 | Bytedance | Lip sync | from 1020 | Premium prompt-directed talking avatar from a still image + audio. 720p / 1080p, up to 60s. People, pets, anime. |
| InfiniTalk | InfiniTalk | Lip sync | from 110 | Audio-driven talking-head from a still image. 480p / 720p. |
| Sync Lipsync v3 | Sync | Lip sync | from 1000 | Dub existing footage — re-syncs lips to a new audio track. Video input, billed per second. |
| Volcengine Lip Sync | Volcengine | Lip sync | from 300 | Video-to-video AI dubbing — re-syncs lips to a new vocal track. Multi-speaker (scene detection + speaker ID) in basic mode. Video input, billed per second. |
The Provider list also offers HeyGen Lipsync Precision and Lipsync 2 Pro for dubbing, and LatentSync, which is best for singing, Wav2Lip, the fastest and cheapest, Video-Retalking, with built-in face enhancement, and SadTalker. Hailuo 3 (minimax-h3) is available through the API, not in the Provider list.
Audio length
| Model | Longest audio |
|---|---|
| Kling Avatar, Kling Avatar Pro | 5 minutes |
| Volcengine Lip Sync | 5 minutes |
| HeyGen Lipsync Precision, Lipsync 2 Pro, Sync Lipsync v3 | 5 minutes, as the reservation limit |
| OmniHuman 1.5 | 60 seconds |
| InfiniTalk | 15 seconds |
| Seedance 2 Fast | 15.2 seconds per clip |
| Seedance 2.5 | 30 seconds per clip |
On Kling Avatar, InfiniTalk, OmniHuman 1.5 and Volcengine Lip Sync, audio longer than the limit is trimmed before the run. On HeyGen Lipsync Precision, Lipsync 2 Pro and Sync Lipsync v3, 5 minutes is the credit reservation limit, not a trim: longer clips reserve at the 5-minute tier. Long runs on Kling Avatar and Volcengine Lip Sync can take tens of minutes, and the editor waits up to about an hour.
Credits
Most models bill per second, rounded up to the next length tier. The credit chip on the node updates when the audio is connected, because the node measures its length. When the length is unknown, the 5-minute tier is reserved. Credits beyond the real cost are refunded when the video is ready.
| Model | 15 s | 30 s | 1 min | 2 min | 5 min |
|---|---|---|---|---|---|
| Kling Avatar (720p) | 300 | 600 | 1,200 | 2,400 | 6,000 |
| Kling Avatar Pro (1080p) | 600 | 1,200 | 2,400 | 4,800 | 12,000 |
| OmniHuman 1.5 | 1,020 | 2,030 | 4,050 | — | — |
| Volcengine Lip Sync | 300 | 600 | 1,200 | 2,400 | 6,000 |
| HeyGen Lipsync Precision | 510 | 1,010 | 2,010 | 4,010 | 10,010 |
| Lipsync 2 Pro | 630 | 1,250 | 2,500 | 5,000 | 12,490 |
| Sync Lipsync v3 | 1,000 | 2,000 | 4,000 | 8,000 | 20,000 |
- OmniHuman 1.5 takes at most 60 seconds of audio, so only the 15-second, 30-second and 1-minute tiers apply. Its resolution does not change the price.
- InfiniTalk has a flat price per run: 110 credits at 480p and 420 at 720p.
- The Seedance models are priced per second, the same way as in Generate Video.
Tips
- Use a clear, front-facing portrait. Face quality matters more than resolution: a sharp 720p photo works better than a blurry 4K one.
- Clean the audio first. Speech without music or noise syncs best. Use Voice Extractor to isolate the voice.
- Add small motions. A motion prompt such as "slight head nods" adds realism.
- Dub in several languages. Translate the speech with Dubbing or a new voiceover, then re-sync the lips with a dubbing model.
- Short clips are cheap. Per-second pricing means a short line costs little; watch the credit chip after you connect the audio.
From the API
When you call Sync Lipsync v3 or Volcengine Lip Sync from the API, SDK or MCP, pass audioDurationSec, the length of the output in seconds. Without it, the run is billed at the 5-minute tier with no refund: 20,000 credits on Sync Lipsync v3 and 6,000 on Volcengine Lip Sync. The editor measures the audio for you. See Run a single node.
Frequently asked questions
Related
AI Avatar
Text to Speech
Dubbing
Voice Extractor
Video models
Last updated on
Cinematic Avatar
Make a short cinematic clip with 1 to 3 HeyGen avatar looks from a text prompt. No script or voice; the prompt directs the scene, action and camera.
Speech to Video
Generate a video driven by a speech track with Wan 2.2. Connect speech audio, a portrait and a prompt, and get a talking video from 30 credits.