# AI Avatar

> Make a talking avatar video with HeyGen. Pick an avatar or animate your own image, then give it a script and a voice, or a recorded voiceover.

Source: https://nodaro.ai/docs/nodes/video/ai-avatar

The **AI Avatar** node makes a talking avatar video with HeyGen. You choose who speaks, a HeyGen avatar look or your own image, and what it says, a script read by a chosen voice or a recorded voiceover. The node returns a video of the avatar speaking, at up to 4K.

- Found in: Video › Animate & Perform
- Output: video
- API type: `ai-avatar`

## When to use it
- A product demo or an explainer with the same presenter in every video.
- Personalized video messages at scale, with one script per recipient from a list.
- A spokesperson video in several languages: change the script and pick a voice in the matching language.
- A news-style or talking-head video from a script another node writes.

For a cinematic scene with avatar looks and no script, use [Cinematic Avatar](https://nodaro.ai/docs/nodes/video/cinematic-avatar). To make any face speak an audio track, use [Lip Sync](https://nodaro.ai/docs/nodes/video/lip-sync).

## Quick start
### Add the node

Press Tab on the canvas and choose **Video › Animate & Perform › AI Avatar**.

### Pick who speaks

On the node card, under **Start with an avatar**, click one of the featured looks, search the catalog, or click **Browse all** to open the full catalog. To animate your own portrait instead, click **Use an image instead** and upload an image, paste a URL, or connect an image node.

### Set the voice and the script

Click the voice on the card to choose one, with a preview for each. Type the script on the card, or connect a text node to the **Script** input. The card shows the number of characters and an estimated duration.

### Run it

When the status bar says the node is ready, click **Run**. The video appears on the node.

Workflow: A written pitch becomes a talking avatar video, which is then joined with product B-roll.

- Text → AI Avatar (script)
- AI Avatar → Combine Videos
- Generate Video → Combine Videos

## Who speaks and how

**Source** decides where the picture comes from:

| Source | What you give | Notes |
| --- | --- | --- |
| **Avatar** (default) | A HeyGen avatar look, chosen from the catalog | Animated by the Avatar IV or Avatar V engine. |
| **Image** | Your own image: connected to the **Image** input, pasted as a URL, or uploaded | No avatar creation or training needed. Image mode uses its own engine, billed like Avatar IV. |

**Speech Mode** decides where the voice comes from:

| Speech Mode | What you give | Voice |
| --- | --- | --- |
| **Text (TTS)** (default) | A script of up to 5,000 characters and a voice from the voice picker | HeyGen text-to-speech in the chosen voice, at the chosen speed. |
| **Wired Audio** | An audio node connected to the **Audio** input | Exactly as recorded, with no text-to-speech. Audio longer than 10 minutes (600 seconds) is trimmed, and the result says so. |

Both sources support both speech modes, and the same voice, background, caption and motion settings.

## The node card

The card itself walks you through the setup, so you rarely need the settings panel for a first run.

| State | What the card shows |
| --- | --- |
| **Empty** | **Start with an avatar**: a row of featured looks, a search box over the whole catalog, **Browse all**, and **Use an image instead**. In image mode, **Start with an image** shows an upload zone and **Choose an avatar** to go back. |
| **Configured** | The portrait with its engine badge and **Change avatar** or **Replace image**; the voice with a preview button; and the script, editable in place, with a character count and an estimated duration. A connected script is shown read-only, with the name of its node. In Wired Audio mode, the card shows the audio connection instead. |
| **Generated** | The video. **Run** makes another version on top of the earlier ones. **New run** hides the results and brings the setup back so you can start fresh; nothing runs until you click Run, and a second click on New run restores the results. |
| **Failed** | The error in red in the status bar. An earlier version stays on show, with a banner that names the failure. |

A status bar at the bottom always says whether the node can run and what is missing, for example "Needs a voice before it can run", and which engine and resolution will render. Text mode needs a script and a voice. Wired Audio needs connected audio. Avatar mode needs an avatar, and image mode needs an image. A connected input counts even before it has produced anything.

## Choose an avatar

**Browse all** opens the **Choose an avatar** window, organized by presenter rather than by look:

- **Search** by name, look or scene. Ctrl+K focuses the search box.
- **Libraries** on the left: **All avatars**, **Your own looks** when the HeyGen account has looks of its own, and **Recently used**. Filter by **Gender** and **Scene**, and switch on the **Avatar V** filter to see only presenters with Avatar V looks.
- **Cards** show one presenter each, with the number of looks. **Load 24 more** shows the next page.
- **The detail column** shows the selected look large, the presenter's other looks, the engines, orientation and default voice, **Preview voice**, and **Use this avatar**.

Picking a look also sets its default voice, when you have not chosen one, and the aspect ratio that matches the look's orientation. A look that HeyGen is still building shows **Processing…** and cannot be picked until it is ready.

The catalog holds thousands of looks and loads in the background; you can pick as soon as you see what you want.

## Settings
| Setting | What it does |
| --- | --- |
| **Source** | **Avatar** or **Image**. See [Who speaks and how](#who-speaks-and-how). |
| **Speech Mode** | **Text (TTS)** or **Wired Audio**. |
| **Avatar** | The HeyGen avatar look, in avatar mode. The picker filters by gender, **Stock** or **Custom** looks, and Avatar V support. |
| **Source Image** | The image to animate, in image mode: connect it to the **Image** input, paste a URL, or upload it. |
| **Voice** | The voice for the script, in text mode, with language, accent and gender filters and a preview. |
| **Script** | What the avatar says, up to 5,000 characters, in text mode. |
| **Voice Speed** | The speaking rate, from 0.5 to 1.5. The default is 1.0. |
| **Engine** | **HeyGen Avatar IV** (the default) or **HeyGen Avatar V**, the premium engine. Avatar mode only. If the look does not support Avatar V, the panel warns you and the node falls back to Avatar IV. |
| **Resolution** | **720p** (the default), **1080p** or **4K**. |
| **Aspect Ratio** | **16:9** (the default) or **9:16**. |
| **Generate captions (SRT)** | Makes a captions file with the video. |

### Advanced settings

Click **Show Advanced** for more control.

| Setting | What it does |
| --- | --- |
| **Pitch** and **Volume** | Adjust the voice. Pitch goes from -50 to +50. Text mode only. |
| **Locale (optional)** | A language locale for the voice, for example en-US. Text mode only. |
| **TTS Engine** | **HeyGen default**, **ElevenLabs** or **Fish**. With ElevenLabs, choose a **Model** (**v3 (recommended)**, Multilingual v2, Turbo v2.5 or Flash v2.5) and tune **Stability**, **Similarity**, **Style** and **Speaker boost**. With Fish, tune **Stability** and **Similarity**. Text mode only. |
| **Background** | **None**, **Color** or **Image** behind the avatar. |
| **Remove background** | Removes the avatar's own background. |
| **Burn captions into video** | Writes the captions into the picture. It also turns on caption generation. |
| **Output Format** | **MP4** or **WebM**. |
| **Fit** | **Cover** or **Contain**. |
| **Motion Prompt (optional)** and **Expressiveness** | Describe the avatar's motion, and choose **Low** (the default), **Medium** or **High** expressiveness. Available with Avatar IV and in image mode. |

## Credits
The price depends on the engine, the resolution and the real length of the video. Avatar V costs more than Avatar IV, and higher resolutions cost more. Image mode is billed like Avatar IV. Captions add no cost. The editor shows the cost before you run.

- **A hold, then the real charge.** Credits are reserved when the run starts. When the video is ready, you pay for its actual length and the rest is refunded.
- **Text mode** reserves from the length of the script and the voice speed. A slower voice reserves a little more, because the video will be longer.
- **Wired Audio** reserves from the measured length of your clip, up to the 10-minute limit.

## When HeyGen is not set up

On a self-hosted install without a HeyGen key, the avatar and voice pickers are empty and show a notice to add a HeyGen key or connect to Nodaro Cloud. A run then fails with the error `heygen_not_configured`. A workflow made elsewhere still shows its avatar and voice. See [Cloud connect](https://nodaro.ai/docs/self-hosting/cloud-connect).

## Tips
- **Draft cheap.** Start with Avatar IV at 720p, then move to 1080p or Avatar V for the final video.
- **Preview voices first.** Listen before you commit a long script to a voice.
- **Use Wired Audio for exact pacing.** A recorded or generated voiceover, for example from [Text to Speech](https://nodaro.ai/docs/nodes/audio/text-to-speech), gives you full control of timing and voice.
- **Go vertical for social.** 9:16 suits TikTok, Instagram Reels and YouTube Shorts.

## Frequently asked questions

### Do I need to create my own avatar?

No. Pick one of the HeyGen avatar looks in the catalog, or choose Use an image instead to animate your own portrait. Image mode needs no avatar creation or training.

### How much does an AI Avatar video cost?

The price depends on the engine, the resolution and the real length of the video. Credits are reserved when the run starts, and the unused part is refunded when the video is ready, so you pay only for the seconds you get.

### What is the difference between Avatar IV and Avatar V?

Both are HeyGen avatar engines. Avatar IV is the default. Avatar V is the premium engine and costs more; a look that does not support it falls back to Avatar IV. Image mode uses its own engine, billed like Avatar IV.

### Can the avatar speak my own recording?

Yes. Set Speech Mode to Wired Audio and connect an audio node. The avatar speaks the recording exactly as it is, with no text-to-speech. Audio longer than 10 minutes is trimmed to 10 minutes.

### What is the difference between AI Avatar and Cinematic Avatar?

AI Avatar makes a talking head that says a script or a voiceover. Cinematic Avatar makes a short cinematic scene from a prompt with 1 to 3 avatar looks, with no script and no voice.
