# Voice Changer Pro

> Detect each speaker in a recording or talking video and give each one a new voice, with the original emotion and timing kept. Priced per minute of speech.

Source: https://nodaro.ai/docs/nodes/audio/voice-changer-pro

The **Voice Changer Pro** node gives every speaker in a conversation a new voice of their own. It detects each speaker in a recording or a talking video, recasts speaker 1 with your first voice, speaker 2 with your second voice, and so on. Each speaker keeps their original emotion, pacing and timing, so a whole dialogue scene, podcast or interview is re-cast in one run.

Voice Changer Pro runs on Nodaro Cloud. On a self-hosted install, the node shows a **NODARO** mark and runs through your [Nodaro Cloud connection](https://nodaro.ai/docs/self-hosting/cloud-connect), billed to the connected nodaro.ai account.

- Found in: Audio › Voices
- Output: audio
- API type: `voice-changer-pro`

## When to use it
- Re-cast a dialogue video with several characters in one step.
- Replace every voice in a podcast or an interview with anonymous voices.
- Cast a scene with specific character voices, and keep the original performance.
- Turn a rough recording with several speakers into a polished voiceover.
- Recast the guests of a panel while the host keeps their own voice.

For a recording with one speaker, [Voice Changer](https://nodaro.ai/docs/nodes/audio/voice-changer) is enough.

## Quick start
### Add the node

Press Tab on the canvas and choose **Audio › Voices › Voice Changer Pro**.

### Wire the recording

Wire an audio node into the **Audio** input, or a talking video into the **Video** input.

### Add one voice per speaker

Open the settings panel. Use **Add a voice** once for each speaker, in the order the speakers first talk. For a speaker who should keep their voice, click **Keep original — don't recast this speaker** instead.

### Run it

Click **Run** on the node. For a video, the re-voiced video appears on the **Video** output and the new sound track on the **Audio** output.

Workflow: An interview with three speakers: the host keeps their voice, the two guests are recast, and the result gets captions.

- Upload Video → Voice Changer Pro (video)
- Voice Changer Pro → Add Captions (video)

## Inputs and outputs
| Input | Accepts | What it does |
| --- | --- | --- |
| **Audio** | Audio nodes, such as Upload Audio and Text to Dialogue | The recording whose speakers are recast. The node returns audio. |
| **Video** | Video nodes, such as Upload Video and Generate Video | The talking video to re-voice. The node returns the video with the new voices. |

When both inputs are wired, the video wins and the audio input is ignored.

| Output | What it carries |
| --- | --- |
| **Audio** | The re-voiced audio. Always produced. For a video, this is the new sound track. |
| **Video** | The re-voiced video. Only produced when a video is wired in. |

**The file decides the mode, not the input.** Before any work starts, the node checks what the file really contains:

- An audio-only file on the **Video** input, such as a voice recording saved as `.mp4` or `.m4a`, runs as audio. You get the re-voiced audio, the **Video** output stays empty, and nothing fails.
- A video on the **Audio** input has its sound extracted first and is treated as audio.

## How speakers get their voices

The node finds each speaker in the recording and numbers them in the order they first talk. Your list of voices is matched to that order.

| Your list | Speaker 1 | Speaker 2 | Speaker 3 | Speaker 4 |
| --- | --- | --- | --- | --- |
| Rachel, Keep original, Aria | Rachel | Original voice | Aria | Original voice |

- **Up to 8 voices.** The list holds 1 to 8 entries, and at least one entry must be a real voice.
- **Keep-slots.** A **Keep original** entry leaves that speaker's voice unchanged, while later speakers are still recast.
- **Shorter lists.** Speakers past the end of the list keep their original voices.
- **Edit the list.** Move a voice with the up and down arrows, turn it into a keep-slot with **Keep**, or remove it.

If you are not sure who talks first, run the recording through [Transcribe](https://nodaro.ai/docs/nodes/audio/transcribe) with **Speaker Diarization** on, and read the speaker labels.

## Settings
| Setting | What it does |
| --- | --- |
| **Add a voice** | Opens the voice browser and adds the chosen voice to the end of the list. Premade, Voice Library and your own voices all work. |
| **Keep original — don't recast this speaker** | Adds a keep-slot to the end of the list. |
| **Model** | **Multilingual v2** (the default) covers 29 languages and is ElevenLabs' recommended model, also for English. **English v2** is for English only. |
| **Separation quality** | **Fast — preserves more of the voice** (the default) or **Quality — finer separation**. |
| **Preserve background music** | **On** (the default) mixes the separated music and sound effects back under the new voices. **Off** drops them for a voice-only result. |
| **Music volume** | The level of the preserved music: **Match source** (the default) keeps its original level, **Normalize** levels its loudness, and **Manual** sets **Music level** from 0 to 200%. Only when Preserve background music is on. |
| **Voice FX** | **None** (the default), or a reverb or echo on the new voices. See [Voice FX](#add-a-room-or-an-echo-to-the-voices). |
| **Remove background noise** | Off by default. Denoises the result. The voices are already isolated automatically, so this option may be unnecessary. |

### Settings for each voice

Open **Voice settings** under a voice to tune that speaker alone. Leave a setting untouched to use the model's default.

| Setting | What it does |
| --- | --- |
| **Conversion Engine** | **Recast** (the default) converts the speech and keeps the original delivery. **Re-speak (v3)** generates a new performance from the transcript with ElevenLabs v3, and understands audio tags. The original delivery is replaced, and on a video the lips no longer match. |
| **Stability** | From 0 to 1 on Recast. On Re-speak: **Most Variable (0)**, **Balanced (0.5)** or **Most Stable (1.0)**. Higher is steadier, lower is more expressive. |
| **Similarity** | From 0 to 1. How closely the result follows the target voice's timbre. Recast only. |
| **Style Exaggeration** | From 0 to 1. Above 0, it amplifies the delivery, with more latency and less stability. Recast only. |
| **Speaker Boost** | Sharpens the likeness to the target voice, with slightly more latency. Recast only. |
| **Volume** | **Match source** (the default) matches the original speaker's loudness. **Normalize** levels the loudness. **Manual** sets a level from 0 to 200%. |
| **Seed** | A whole number from 0 to 4,294,967,295. The same source, settings and seed recast this speaker identically. Leave it empty for a random seed. |

## How the background is kept

Before any voice is recast, the node always splits the recording into a voice track and a music-and-effects track. **Preserve background music** does not decide whether this split happens. It only decides whether the music track is mixed back under the new voices.

If music bleeds into the new voices, or a voice sounds thin, set **Separation quality** to **Quality — finer separation**. Keep **Fast** when speed matters or the voices already sound clean.

## Add a room or an echo to the voices

**Voice FX** adds one effect to all the new voices together, before the music is mixed back in. The effect sits on the voices only, so with **Preserve background music** on, the music stays dry under a voice with reverb.

- **Reverb spaces:** Room (indoor), Bathroom, Car interior, Hall / Lobby, Concert Hall, Church, Cave, Arena / Stadium and Outdoor (open air). **Wet / Dry mix** from 0 to 100 sets how much room you hear.
- **Character:** Telephone, Megaphone / PA, Echo / Slap-back and Custom. Echo and Custom use **Delay (ms)**, from 20 to 2,000, and **Decay**, from 0 to 1, where a higher decay gives more repeats.

The [Audio FX](https://nodaro.ai/docs/nodes/audio/audio-fx) page describes each preset.

## Re-voice a whole video

Wire any talking video into the **Video** input. The node then:

1. Extracts the sound track from the clip.
2. Detects each speaker in the order they first talk.
3. Recasts each speaker with the matching voice.
4. Puts the new voices back on the original video, and returns the video and the new sound track.

**The video needs a sound track.** Most text-to-video and image-to-video models make silent video, and a silent clip fails at once. Use a clip with spoken audio, or wire a recording into the **Audio** input.

## Credits
Only recast voices cost credits. Keep-slots and speakers past the end of your list pass through for free.

- **A Recast voice** is billed by the length of its stem. The stem runs from the start of the clip to that speaker's last line. The rate is 40 credits per minute, counted per second and rounded up to the next credit.
- **A Re-speak voice** is billed at 30 credits per started 1,000 characters of the text it speaks.
- **Minimum.** Every voice costs at least 4 credits, which is 6 seconds of speech, and so does every run.

| Recast voices | Credits |
| --- | --- |
| 1 Recast voice whose last line ends at 26.76 seconds | 18 |
| 1 Recast voice with a 60-second stem | 40 |
| 1 Recast voice with a 61-second stem | 41 |
| 1 Recast voice with a 3-second stem | 4 (the minimum) |
| 2 Recast voices with 60-second stems, and 1 Re-speak voice of 1,500 characters | 40 + 40 + 60 = 140 |

You are charged for the stems that were actually converted, and never more than the amount held when the run started.

## Use it as an app

Voice Changer Pro is also a standalone app at [voice.nodaro.ai](https://voice.nodaro.ai). Paste a link or upload a video, choose a voice for each speaker, and get the converted video back.

## Tips
- **Order the voices with care.** The match is by position: the first voice goes to the first speaker to talk.
- **Tune each speaker separately.** Each voice has its own settings, so you can steady one speaker and keep another expressive.
- **Balance loudness per speaker.** Use **Volume** on each voice: **Match source** mirrors the original speaker, and **Manual** sets an exact level.
- **Map fewer voices than speakers.** Only the listed speakers change; everyone after them keeps their voice.
- **Keep the performance.** Recast converts speech to speech, so the original acting carries through. Use **Re-speak (v3)** only when you want a new performance, for example on an audio-only project.

## Troubleshooting
**One speaker sounds robotic.** Lower that voice's **Similarity** or **Stability**, or choose a voice whose timbre is closer to the original speaker.

**Music bleeds into the voices.** Set **Separation quality** to **Quality — finer separation**.

**The wrong speaker got a voice.** The speakers are numbered by who talks first. Check the order with [Transcribe](https://nodaro.ai/docs/nodes/audio/transcribe) and reorder the voices.

**A self-hosted run is refused with `nodaro_connection_required`.** The install has no Nodaro Cloud connection. Connect it with OAuth or an API key, as described in [Cloud connect](https://nodaro.ai/docs/self-hosting/cloud-connect).

## From the API
`POST /v1/voice-changer-pro` runs the node from code, on Nodaro Cloud. `orderedVoices` maps speaker N to entry N. An entry is a voice id, an object with per-voice settings, or `null` for a keep-slot. The settings object takes `voiceId`, `engine` (`"sts"` or `"v3"`), `stability`, `similarityBoost`, `style`, `useSpeakerBoost`, `seed`, `volumeMode` and `volume`. With `engine: "v3"`, `stability` accepts only 0, 0.5 or 1, and `similarityBoost`, `style` and `useSpeakerBoost` are ignored. The result's `sourceHasVideo` tells you whether the run used video mode or audio mode.

On Nodaro Cloud, you can also run the same engine in three steps. You can then check the speakers before you pay for a recast, and mix the tracks before you render:

1. **Analyze.** `POST /v1/voice-changer-pro/analyze` separates the voices from the music and detects the speakers. It returns each speaker's time segments, word count and a transcript snippet, plus the detected language. It costs 10 credits.
2. **Recast to stems.** Call `POST /v1/voice-changer-pro` with `output: "stems"` and the analysis. The speakers are not detected again, so you can recast with other voices as often as you like. The result is one dry track per voice.
3. **Export.** `POST /v1/voice-changer-pro/export` renders the final video from your mix. Send up to 16 tracks, each with a `gain` from 0 to 200, `muted`, and a `kind` of `voice` or `background`, plus an optional `voiceFx`. At least one track must be unmuted. The video stream is copied, never re-encoded. Adjusting the mix is free; the export costs 1 credit.

The SDK methods are `client.voices.analyze`, `client.voices.recast` and `client.voices.exportMix`. The CLI commands are `nodaro voice analyze`, `nodaro voice recast` and `nodaro voice export`. The MCP tools are `voice_changer_pro`, `voice_changer_pro_analyze` and `voice_changer_pro_export`. See [Voice and media](https://nodaro.ai/docs/developers/api/voice-and-media).

## Frequently asked questions

### How does Voice Changer Pro know which voice goes to which speaker?

It detects the speakers in the order they first talk. Voice 1 recasts the first speaker to talk, voice 2 the second, and so on. Speakers past the end of your list keep their original voices.

### Can I change only some of the speakers?

Yes. Add a Keep original slot for each speaker who should keep their own voice. For example, Rachel, Keep original, Aria recasts speakers 1 and 3, and speaker 2 is unchanged. Kept speakers cost nothing.

### How many credits does Voice Changer Pro cost?

Each recast voice costs 40 credits per minute of its stem, rounded up to the next credit. The stem runs from the start of the clip to that speaker's last line. A Re-speak voice costs 30 credits per started 1,000 characters, and every voice and every run costs at least 4 credits.

### Does Voice Changer Pro keep the background music?

Yes, by default. The music and sound effects are separated from the voices first, and Preserve background music mixes them back under the new voices. Turn it off for a voice-only result.

### Can I use Voice Changer Pro on a self-hosted install?

Yes, through a Nodaro Cloud connection. The node runs on Nodaro Cloud and is billed to the connected nodaro.ai account. Without a connection, the node shows a Connect nodaro.ai button.
