Nodaro Docs
DocumentationNode ReferenceModelsAI Agents (MCP)DevelopersSelf-hostingResearch
Video

AI Avatar

Make a talking avatar video with HeyGen. Pick an avatar or animate your own image, then give it a script and a voice, or a recorded voiceover.

The AI Avatar node makes a talking avatar video with HeyGen. You choose who speaks, a HeyGen avatar look or your own image, and what it says, a script read by a chosen voice or a recorded voiceover. The node returns a video of the avatar speaking, at up to 4K.

When to use it

  • A product demo or an explainer with the same presenter in every video.
  • Personalized video messages at scale, with one script per recipient from a list.
  • A spokesperson video in several languages: change the script and pick a voice in the matching language.
  • A news-style or talking-head video from a script another node writes.

For a cinematic scene with avatar looks and no script, use Cinematic Avatar. To make any face speak an audio track, use Lip Sync.

Quick start

Add the node

Press Tab on the canvas and choose Video › Animate & Perform › AI Avatar.

Pick who speaks

On the node card, under Start with an avatar, click one of the featured looks, search the catalog, or click Browse all to open the full catalog. To animate your own portrait instead, click Use an image instead and upload an image, paste a URL, or connect an image node.

Set the voice and the script

Click the voice on the card to choose one, with a preview for each. Type the script on the card, or connect a text node to the Script input. The card shows the number of characters and an estimated duration.

Run it

When the status bar says the node is ready, click Run. The video appears on the node.

scriptTextProduct pitchAI AvatarAvatar IV · 9:16Generate VideoProduct B-rollCombine Videos
A written pitch becomes a talking avatar video, which is then joined with product B-roll.

Who speaks and how

Source decides where the picture comes from:

SourceWhat you giveNotes
Avatar (default)A HeyGen avatar look, chosen from the catalogAnimated by the Avatar IV or Avatar V engine.
ImageYour own image: connected to the Image input, pasted as a URL, or uploadedNo avatar creation or training needed. Image mode uses its own engine, billed like Avatar IV.

Speech Mode decides where the voice comes from:

Speech ModeWhat you giveVoice
Text (TTS) (default)A script of up to 5,000 characters and a voice from the voice pickerHeyGen text-to-speech in the chosen voice, at the chosen speed.
Wired AudioAn audio node connected to the Audio inputExactly as recorded, with no text-to-speech. Audio longer than 10 minutes (600 seconds) is trimmed, and the result says so.

Both sources support both speech modes, and the same voice, background, caption and motion settings.

The node card

The card itself walks you through the setup, so you rarely need the settings panel for a first run.

StateWhat the card shows
EmptyStart with an avatar: a row of featured looks, a search box over the whole catalog, Browse all, and Use an image instead. In image mode, Start with an image shows an upload zone and Choose an avatar to go back.
ConfiguredThe portrait with its engine badge and Change avatar or Replace image; the voice with a preview button; and the script, editable in place, with a character count and an estimated duration. A connected script is shown read-only, with the name of its node. In Wired Audio mode, the card shows the audio connection instead.
GeneratedThe video. Run makes another version on top of the earlier ones. New run hides the results and brings the setup back so you can start fresh; nothing runs until you click Run, and a second click on New run restores the results.
FailedThe error in red in the status bar. An earlier version stays on show, with a banner that names the failure.

A status bar at the bottom always says whether the node can run and what is missing, for example "Needs a voice before it can run", and which engine and resolution will render. Text mode needs a script and a voice. Wired Audio needs connected audio. Avatar mode needs an avatar, and image mode needs an image. A connected input counts even before it has produced anything.

Choose an avatar

Browse all opens the Choose an avatar window, organized by presenter rather than by look:

  • Search by name, look or scene. Ctrl+K focuses the search box.
  • Libraries on the left: All avatars, Your own looks when the HeyGen account has looks of its own, and Recently used. Filter by Gender and Scene, and switch on the Avatar V filter to see only presenters with Avatar V looks.
  • Cards show one presenter each, with the number of looks. Load 24 more shows the next page.
  • The detail column shows the selected look large, the presenter's other looks, the engines, orientation and default voice, Preview voice, and Use this avatar.

Picking a look also sets its default voice, when you have not chosen one, and the aspect ratio that matches the look's orientation. A look that HeyGen is still building shows Processing… and cannot be picked until it is ready.

The catalog holds thousands of looks and loads in the background; you can pick as soon as you see what you want.

Settings

SettingWhat it does
SourceAvatar or Image. See Who speaks and how.
Speech ModeText (TTS) or Wired Audio.
AvatarThe HeyGen avatar look, in avatar mode. The picker filters by gender, Stock or Custom looks, and Avatar V support.
Source ImageThe image to animate, in image mode: connect it to the Image input, paste a URL, or upload it.
VoiceThe voice for the script, in text mode, with language, accent and gender filters and a preview.
ScriptWhat the avatar says, up to 5,000 characters, in text mode.
Voice SpeedThe speaking rate, from 0.5 to 1.5. The default is 1.0.
EngineHeyGen Avatar IV (the default) or HeyGen Avatar V, the premium engine. Avatar mode only. If the look does not support Avatar V, the panel warns you and the node falls back to Avatar IV.
Resolution720p (the default), 1080p or 4K.
Aspect Ratio16:9 (the default) or 9:16.
Generate captions (SRT)Makes a captions file with the video.

Advanced settings

Click Show Advanced for more control.

SettingWhat it does
Pitch and VolumeAdjust the voice. Pitch goes from -50 to +50. Text mode only.
Locale (optional)A language locale for the voice, for example en-US. Text mode only.
TTS EngineHeyGen default, ElevenLabs or Fish. With ElevenLabs, choose a Model (v3 (recommended), Multilingual v2, Turbo v2.5 or Flash v2.5) and tune Stability, Similarity, Style and Speaker boost. With Fish, tune Stability and Similarity. Text mode only.
BackgroundNone, Color or Image behind the avatar.
Remove backgroundRemoves the avatar's own background.
Burn captions into videoWrites the captions into the picture. It also turns on caption generation.
Output FormatMP4 or WebM.
FitCover or Contain.
Motion Prompt (optional) and ExpressivenessDescribe the avatar's motion, and choose Low (the default), Medium or High expressiveness. Available with Avatar IV and in image mode.

Credits

The price depends on the engine, the resolution and the real length of the video. Avatar V costs more than Avatar IV, and higher resolutions cost more. Image mode is billed like Avatar IV. Captions add no cost. The editor shows the cost before you run.

  • A hold, then the real charge. Credits are reserved when the run starts. When the video is ready, you pay for its actual length and the rest is refunded.
  • Text mode reserves from the length of the script and the voice speed. A slower voice reserves a little more, because the video will be longer.
  • Wired Audio reserves from the measured length of your clip, up to the 10-minute limit.

When HeyGen is not set up

On a self-hosted install without a HeyGen key, the avatar and voice pickers are empty and show a notice to add a HeyGen key or connect to Nodaro Cloud. A run then fails with the error heygen_not_configured. A workflow made elsewhere still shows its avatar and voice. See Cloud connect.

Tips

  • Draft cheap. Start with Avatar IV at 720p, then move to 1080p or Avatar V for the final video.
  • Preview voices first. Listen before you commit a long script to a voice.
  • Use Wired Audio for exact pacing. A recorded or generated voiceover, for example from Text to Speech, gives you full control of timing and voice.
  • Go vertical for social. 9:16 suits TikTok, Instagram Reels and YouTube Shorts.

Frequently asked questions

Last updated on

On this page