AI Avatar
Make a talking avatar video with HeyGen. Pick an avatar or animate your own image, then give it a script and a voice, or a recorded voiceover.
The AI Avatar node makes a talking avatar video with HeyGen. You choose who speaks, a HeyGen avatar look or your own image, and what it says, a script read by a chosen voice or a recorded voiceover. The node returns a video of the avatar speaking, at up to 4K.
When to use it
- A product demo or an explainer with the same presenter in every video.
- Personalized video messages at scale, with one script per recipient from a list.
- A spokesperson video in several languages: change the script and pick a voice in the matching language.
- A news-style or talking-head video from a script another node writes.
For a cinematic scene with avatar looks and no script, use Cinematic Avatar. To make any face speak an audio track, use Lip Sync.
Quick start
Add the node
Press Tab on the canvas and choose Video › Animate & Perform › AI Avatar.
Pick who speaks
On the node card, under Start with an avatar, click one of the featured looks, search the catalog, or click Browse all to open the full catalog. To animate your own portrait instead, click Use an image instead and upload an image, paste a URL, or connect an image node.
Set the voice and the script
Click the voice on the card to choose one, with a preview for each. Type the script on the card, or connect a text node to the Script input. The card shows the number of characters and an estimated duration.
Run it
When the status bar says the node is ready, click Run. The video appears on the node.
Who speaks and how
Source decides where the picture comes from:
| Source | What you give | Notes |
|---|---|---|
| Avatar (default) | A HeyGen avatar look, chosen from the catalog | Animated by the Avatar IV or Avatar V engine. |
| Image | Your own image: connected to the Image input, pasted as a URL, or uploaded | No avatar creation or training needed. Image mode uses its own engine, billed like Avatar IV. |
Speech Mode decides where the voice comes from:
| Speech Mode | What you give | Voice |
|---|---|---|
| Text (TTS) (default) | A script of up to 5,000 characters and a voice from the voice picker | HeyGen text-to-speech in the chosen voice, at the chosen speed. |
| Wired Audio | An audio node connected to the Audio input | Exactly as recorded, with no text-to-speech. Audio longer than 10 minutes (600 seconds) is trimmed, and the result says so. |
Both sources support both speech modes, and the same voice, background, caption and motion settings.
The node card
The card itself walks you through the setup, so you rarely need the settings panel for a first run.
| State | What the card shows |
|---|---|
| Empty | Start with an avatar: a row of featured looks, a search box over the whole catalog, Browse all, and Use an image instead. In image mode, Start with an image shows an upload zone and Choose an avatar to go back. |
| Configured | The portrait with its engine badge and Change avatar or Replace image; the voice with a preview button; and the script, editable in place, with a character count and an estimated duration. A connected script is shown read-only, with the name of its node. In Wired Audio mode, the card shows the audio connection instead. |
| Generated | The video. Run makes another version on top of the earlier ones. New run hides the results and brings the setup back so you can start fresh; nothing runs until you click Run, and a second click on New run restores the results. |
| Failed | The error in red in the status bar. An earlier version stays on show, with a banner that names the failure. |
A status bar at the bottom always says whether the node can run and what is missing, for example "Needs a voice before it can run", and which engine and resolution will render. Text mode needs a script and a voice. Wired Audio needs connected audio. Avatar mode needs an avatar, and image mode needs an image. A connected input counts even before it has produced anything.
Choose an avatar
Browse all opens the Choose an avatar window, organized by presenter rather than by look:
- Search by name, look or scene. Ctrl+K focuses the search box.
- Libraries on the left: All avatars, Your own looks when the HeyGen account has looks of its own, and Recently used. Filter by Gender and Scene, and switch on the Avatar V filter to see only presenters with Avatar V looks.
- Cards show one presenter each, with the number of looks. Load 24 more shows the next page.
- The detail column shows the selected look large, the presenter's other looks, the engines, orientation and default voice, Preview voice, and Use this avatar.
Picking a look also sets its default voice, when you have not chosen one, and the aspect ratio that matches the look's orientation. A look that HeyGen is still building shows Processing… and cannot be picked until it is ready.
The catalog holds thousands of looks and loads in the background; you can pick as soon as you see what you want.
Settings
| Setting | What it does |
|---|---|
| Source | Avatar or Image. See Who speaks and how. |
| Speech Mode | Text (TTS) or Wired Audio. |
| Avatar | The HeyGen avatar look, in avatar mode. The picker filters by gender, Stock or Custom looks, and Avatar V support. |
| Source Image | The image to animate, in image mode: connect it to the Image input, paste a URL, or upload it. |
| Voice | The voice for the script, in text mode, with language, accent and gender filters and a preview. |
| Script | What the avatar says, up to 5,000 characters, in text mode. |
| Voice Speed | The speaking rate, from 0.5 to 1.5. The default is 1.0. |
| Engine | HeyGen Avatar IV (the default) or HeyGen Avatar V, the premium engine. Avatar mode only. If the look does not support Avatar V, the panel warns you and the node falls back to Avatar IV. |
| Resolution | 720p (the default), 1080p or 4K. |
| Aspect Ratio | 16:9 (the default) or 9:16. |
| Generate captions (SRT) | Makes a captions file with the video. |
Advanced settings
Click Show Advanced for more control.
| Setting | What it does |
|---|---|
| Pitch and Volume | Adjust the voice. Pitch goes from -50 to +50. Text mode only. |
| Locale (optional) | A language locale for the voice, for example en-US. Text mode only. |
| TTS Engine | HeyGen default, ElevenLabs or Fish. With ElevenLabs, choose a Model (v3 (recommended), Multilingual v2, Turbo v2.5 or Flash v2.5) and tune Stability, Similarity, Style and Speaker boost. With Fish, tune Stability and Similarity. Text mode only. |
| Background | None, Color or Image behind the avatar. |
| Remove background | Removes the avatar's own background. |
| Burn captions into video | Writes the captions into the picture. It also turns on caption generation. |
| Output Format | MP4 or WebM. |
| Fit | Cover or Contain. |
| Motion Prompt (optional) and Expressiveness | Describe the avatar's motion, and choose Low (the default), Medium or High expressiveness. Available with Avatar IV and in image mode. |
Credits
The price depends on the engine, the resolution and the real length of the video. Avatar V costs more than Avatar IV, and higher resolutions cost more. Image mode is billed like Avatar IV. Captions add no cost. The editor shows the cost before you run.
- A hold, then the real charge. Credits are reserved when the run starts. When the video is ready, you pay for its actual length and the rest is refunded.
- Text mode reserves from the length of the script and the voice speed. A slower voice reserves a little more, because the video will be longer.
- Wired Audio reserves from the measured length of your clip, up to the 10-minute limit.
When HeyGen is not set up
On a self-hosted install without a HeyGen key, the avatar and voice pickers are empty and show a notice to add a HeyGen key or connect to Nodaro Cloud. A run then fails with the error heygen_not_configured. A workflow made elsewhere still shows its avatar and voice. See Cloud connect.
Tips
- Draft cheap. Start with Avatar IV at 720p, then move to 1080p or Avatar V for the final video.
- Preview voices first. Listen before you commit a long script to a voice.
- Use Wired Audio for exact pacing. A recorded or generated voiceover, for example from Text to Speech, gives you full control of timing and voice.
- Go vertical for social. 9:16 suits TikTok, Instagram Reels and YouTube Shorts.
Frequently asked questions
Related
Cinematic Avatar
Lip Sync
Text to Speech
Add Captions
Last updated on
Generate Script
Write a multi-scene video script with AI. Each scene gets a description, action, mood, length, camera, dialogue and a ready-to-use image prompt.
Cinematic Avatar
Make a short cinematic clip with 1 to 3 HeyGen avatar looks from a text prompt. No script or voice; the prompt directs the scene, action and camera.