Nodaro Docs
DocumentationNode ReferenceModelsAI Agents (MCP)DevelopersSelf-hostingResearch
Video

Lip Sync

Make a portrait speak an audio track, or dub an existing video by re-syncing its lips to new speech, with Kling Avatar, OmniHuman, Seedance and more.

The Lip Sync node makes a face speak an audio track. Give it a portrait and a voice track, and it makes a talking-head video. Give it a video and new audio, and it re-syncs the lips in the video to the new speech, which dubs the clip. You choose the model, and the node shows the inputs that model needs.

When to use it

  • A talking head from a single photo, for example an AI-generated character with a voiceover.
  • A spokesperson video for a product demo.
  • A dubbed version of a video: translated speech, re-synced to the speaker's lips.
  • A performance directed by a prompt, such as singing into a microphone, with OmniHuman 1.5.

For a full avatar video from a script, with voices and captions built in, use AI Avatar.

Quick start

Add the node

Press Tab on the canvas and choose Video › Animate & Perform › Lip Sync.

Choose the model

Open the settings panel and choose a model under Provider. The node then shows the inputs that model needs: Portrait, Source video, or both.

Connect the face and the voice

Connect an image to Portrait, or a video to Source video, and an audio node to Audio, for example Text to Speech.

Run it

Click Run. The lip-synced video appears on the node.

portraitaudioGenerate ImageFront-facing portraitTextThe line to sayText to SpeechElevenLabs v3Lip SyncKling Avatar Pro
A portrait and a generated voice line become a talking-head video.

Portrait or video

What you connectWhat the node makesModels
A Portrait and AudioA new talking-head video from a still image.Kling Avatar, Kling Avatar Pro, InfiniTalk, OmniHuman 1.5, Seedance 2, Seedance 2 Fast, Seedance 2 Mini, Seedance 2.5, SadTalker
A Source video and AudioThe same video with its lips re-synced to the new audio.Volcengine Lip Sync, Sync Lipsync v3, Lipsync 2 Pro, HeyGen Lipsync Precision, LatentSync, Video-Retalking
Either one, and AudioWav2Lip accepts a portrait or a video.Wav2Lip

Inputs

InputAcceptsWhat it does
AudioAudio nodesThe speech or voiceover. Required for every model.
PortraitImage nodesA clear face photo, for portrait models.
Source videoVideo nodesThe clip to dub, for dubbing models.

The output, Video, is the URL of the lip-synced video. When several nodes are connected, the node card lets you choose which image, video and audio to use.

Settings

SettingWhat it does
ProviderThe lip-sync model. The default is Kling Avatar.
ResolutionOn Kling Avatar, Kling Avatar Pro, InfiniTalk and the Seedance models: 480p or 720p, with 720p as the default. Seedance 2 also offers 1080p. OmniHuman 1.5 offers 720p or 1080p, with 1080p as the default.
Motion Prompt (optional)On the same models: head and expression motion, for example "slight head nods" or "expressive eyebrows". On OmniHuman 1.5 it directs the whole performance.
Pre & post textText always added before and after the motion prompt. See Prompt pre and post text.

Settings for specific models

ModelSettings
OmniHuman 1.5Motion Prompt directs the performance, for example "sing confidently into a microphone". Fast Mode trades some quality for speed. Seed makes a result repeatable; -1 means random. It animates people, pets and anime at any aspect ratio.
HeyGen Lipsync PrecisionDynamic Duration (on) adjusts the output length to the new audio. Remove Music Track (off) strips background music from the source video. Speech Enhancement (off) improves speech clarity.
Lipsync 2 ProSync Mode sets what happens when the audio and the video have different lengths: Loop (the default), Bounce, Cut off, Silence or Remap. Temperature, from 0 to 1, sets how expressive the lip sync is; the default is 0.5. Active Speaker Detection syncs whoever is speaking.
Sync Lipsync v3Sync Mode, with Cut off as the default, and Active Speaker Detection, which you should turn on for a video with more than one person. The model manages its own expressiveness.
Volcengine Lip SyncMode: Lite for one frontal speaker, faster, or Basic for complex scenes with several speakers. Separate vocals (denoise) strips noise from the audio. Scene detection + speaker ID (Basic only) finds scene cuts and who speaks. Loop video if audio is longer (Lite only, on by default) and Reverse loop (ping-pong) handle audio that runs past the video. Template start time (seconds) sets where in the video the lips start to follow. The output is as long as the audio: a longer video is cut, and a shorter one loops.

The Seedance models do native, phoneme-level lip sync in more than 8 languages. They work like Generate Video with the audio sent as reference audio.

Models

ModelMakerModesCreditsDetails
minimax-h3MiniMaxImage to video, Text to videofrom 230MiniMax Hailuo 3 — premium multimodal tier: first/last frame + image/video/audio references, native audio, 2K (default) or 768P output, 4-15s per-second pricing.
Kling Avatar StandardKuaishouLip sync280Lip-sync a still portrait to driving audio. Standard quality.
Kling Avatar ProKuaishouLip sync560Premium lip-sync — better mouth shape and timing.
Seedance 2BytedanceImage to video, Text to videofrom 230Seedance 2 — premium tier with native audio. Per-second pricing by resolution.
Seedance 2 FastBytedanceImage to video, Text to videofrom 180Cheaper / quicker Seedance 2 tier.
Seedance 2 MiniBytedanceImage to video, Text to videofrom 120Budget Seedance 2 tier — 480p/720p only, per-second pricing by resolution.
Seedance 2.5BytedanceImage to video, Text to videofrom 340Seedance 2.5 — up to 30s in one shot, native audio, wide multimodal references. 480p/720p/1080p.
OmniHuman 1.5BytedanceLip syncfrom 1020Premium prompt-directed talking avatar from a still image + audio. 720p / 1080p, up to 60s. People, pets, anime.
InfiniTalkInfiniTalkLip syncfrom 110Audio-driven talking-head from a still image. 480p / 720p.
Sync Lipsync v3SyncLip syncfrom 1000Dub existing footage — re-syncs lips to a new audio track. Video input, billed per second.
Volcengine Lip SyncVolcengineLip syncfrom 300Video-to-video AI dubbing — re-syncs lips to a new vocal track. Multi-speaker (scene detection + speaker ID) in basic mode. Video input, billed per second.

The Provider list also offers HeyGen Lipsync Precision and Lipsync 2 Pro for dubbing, and LatentSync, which is best for singing, Wav2Lip, the fastest and cheapest, Video-Retalking, with built-in face enhancement, and SadTalker. Hailuo 3 (minimax-h3) is available through the API, not in the Provider list.

Audio length

ModelLongest audio
Kling Avatar, Kling Avatar Pro5 minutes
Volcengine Lip Sync5 minutes
HeyGen Lipsync Precision, Lipsync 2 Pro, Sync Lipsync v35 minutes, as the reservation limit
OmniHuman 1.560 seconds
InfiniTalk15 seconds
Seedance 2 Fast15.2 seconds per clip
Seedance 2.530 seconds per clip

On Kling Avatar, InfiniTalk, OmniHuman 1.5 and Volcengine Lip Sync, audio longer than the limit is trimmed before the run. On HeyGen Lipsync Precision, Lipsync 2 Pro and Sync Lipsync v3, 5 minutes is the credit reservation limit, not a trim: longer clips reserve at the 5-minute tier. Long runs on Kling Avatar and Volcengine Lip Sync can take tens of minutes, and the editor waits up to about an hour.

Credits

Most models bill per second, rounded up to the next length tier. The credit chip on the node updates when the audio is connected, because the node measures its length. When the length is unknown, the 5-minute tier is reserved. Credits beyond the real cost are refunded when the video is ready.

Model15 s30 s1 min2 min5 min
Kling Avatar (720p)3006001,2002,4006,000
Kling Avatar Pro (1080p)6001,2002,4004,80012,000
OmniHuman 1.51,0202,0304,050——
Volcengine Lip Sync3006001,2002,4006,000
HeyGen Lipsync Precision5101,0102,0104,01010,010
Lipsync 2 Pro6301,2502,5005,00012,490
Sync Lipsync v31,0002,0004,0008,00020,000
  • OmniHuman 1.5 takes at most 60 seconds of audio, so only the 15-second, 30-second and 1-minute tiers apply. Its resolution does not change the price.
  • InfiniTalk has a flat price per run: 110 credits at 480p and 420 at 720p.
  • The Seedance models are priced per second, the same way as in Generate Video.

Tips

  • Use a clear, front-facing portrait. Face quality matters more than resolution: a sharp 720p photo works better than a blurry 4K one.
  • Clean the audio first. Speech without music or noise syncs best. Use Voice Extractor to isolate the voice.
  • Add small motions. A motion prompt such as "slight head nods" adds realism.
  • Dub in several languages. Translate the speech with Dubbing or a new voiceover, then re-sync the lips with a dubbing model.
  • Short clips are cheap. Per-second pricing means a short line costs little; watch the credit chip after you connect the audio.

From the API

When you call Sync Lipsync v3 or Volcengine Lip Sync from the API, SDK or MCP, pass audioDurationSec, the length of the output in seconds. Without it, the run is billed at the 5-minute tier with no refund: 20,000 credits on Sync Lipsync v3 and 6,000 on Volcengine Lip Sync. The editor measures the audio for you. See Run a single node.

Frequently asked questions

Last updated on

On this page