Voice Changer Pro
Detect each speaker in a recording or talking video and give each one a new voice, with the original emotion and timing kept. Priced per minute of speech.
Available on Nodaro Cloud
The Voice Changer Pro node gives every speaker in a conversation a new voice of their own. It detects each speaker in a recording or a talking video, recasts speaker 1 with your first voice, speaker 2 with your second voice, and so on. Each speaker keeps their original emotion, pacing and timing, so a whole dialogue scene, podcast or interview is re-cast in one run.
Voice Changer Pro runs on Nodaro Cloud. On a self-hosted install, the node shows a NODARO mark and runs through your Nodaro Cloud connection, billed to the connected nodaro.ai account.
When to use it
- Re-cast a dialogue video with several characters in one step.
- Replace every voice in a podcast or an interview with anonymous voices.
- Cast a scene with specific character voices, and keep the original performance.
- Turn a rough recording with several speakers into a polished voiceover.
- Recast the guests of a panel while the host keeps their own voice.
For a recording with one speaker, Voice Changer is enough.
Quick start
Add the node
Press Tab on the canvas and choose Audio › Voices › Voice Changer Pro.
Wire the recording
Wire an audio node into the Audio input, or a talking video into the Video input.
Add one voice per speaker
Open the settings panel. Use Add a voice once for each speaker, in the order the speakers first talk. For a speaker who should keep their voice, click Keep original — don't recast this speaker instead.
Run it
Click Run on the node. For a video, the re-voiced video appears on the Video output and the new sound track on the Audio output.
Inputs and outputs
| Input | Accepts | What it does |
|---|---|---|
| Audio | Audio nodes, such as Upload Audio and Text to Dialogue | The recording whose speakers are recast. The node returns audio. |
| Video | Video nodes, such as Upload Video and Generate Video | The talking video to re-voice. The node returns the video with the new voices. |
When both inputs are wired, the video wins and the audio input is ignored.
| Output | What it carries |
|---|---|
| Audio | The re-voiced audio. Always produced. For a video, this is the new sound track. |
| Video | The re-voiced video. Only produced when a video is wired in. |
The file decides the mode, not the input. Before any work starts, the node checks what the file really contains:
- An audio-only file on the Video input, such as a voice recording saved as
.mp4or.m4a, runs as audio. You get the re-voiced audio, the Video output stays empty, and nothing fails. - A video on the Audio input has its sound extracted first and is treated as audio.
How speakers get their voices
The node finds each speaker in the recording and numbers them in the order they first talk. Your list of voices is matched to that order.
| Your list | Speaker 1 | Speaker 2 | Speaker 3 | Speaker 4 |
|---|---|---|---|---|
| Rachel, Keep original, Aria | Rachel | Original voice | Aria | Original voice |
- Up to 8 voices. The list holds 1 to 8 entries, and at least one entry must be a real voice.
- Keep-slots. A Keep original entry leaves that speaker's voice unchanged, while later speakers are still recast.
- Shorter lists. Speakers past the end of the list keep their original voices.
- Edit the list. Move a voice with the up and down arrows, turn it into a keep-slot with Keep, or remove it.
If you are not sure who talks first, run the recording through Transcribe with Speaker Diarization on, and read the speaker labels.
Settings
| Setting | What it does |
|---|---|
| Add a voice | Opens the voice browser and adds the chosen voice to the end of the list. Premade, Voice Library and your own voices all work. |
| Keep original — don't recast this speaker | Adds a keep-slot to the end of the list. |
| Model | Multilingual v2 (the default) covers 29 languages and is ElevenLabs' recommended model, also for English. English v2 is for English only. |
| Separation quality | Fast — preserves more of the voice (the default) or Quality — finer separation. |
| Preserve background music | On (the default) mixes the separated music and sound effects back under the new voices. Off drops them for a voice-only result. |
| Music volume | The level of the preserved music: Match source (the default) keeps its original level, Normalize levels its loudness, and Manual sets Music level from 0 to 200%. Only when Preserve background music is on. |
| Voice FX | None (the default), or a reverb or echo on the new voices. See Voice FX. |
| Remove background noise | Off by default. Denoises the result. The voices are already isolated automatically, so this option may be unnecessary. |
Settings for each voice
Open Voice settings under a voice to tune that speaker alone. Leave a setting untouched to use the model's default.
| Setting | What it does |
|---|---|
| Conversion Engine | Recast (the default) converts the speech and keeps the original delivery. Re-speak (v3) generates a new performance from the transcript with ElevenLabs v3, and understands audio tags. The original delivery is replaced, and on a video the lips no longer match. |
| Stability | From 0 to 1 on Recast. On Re-speak: Most Variable (0), Balanced (0.5) or Most Stable (1.0). Higher is steadier, lower is more expressive. |
| Similarity | From 0 to 1. How closely the result follows the target voice's timbre. Recast only. |
| Style Exaggeration | From 0 to 1. Above 0, it amplifies the delivery, with more latency and less stability. Recast only. |
| Speaker Boost | Sharpens the likeness to the target voice, with slightly more latency. Recast only. |
| Volume | Match source (the default) matches the original speaker's loudness. Normalize levels the loudness. Manual sets a level from 0 to 200%. |
| Seed | A whole number from 0 to 4,294,967,295. The same source, settings and seed recast this speaker identically. Leave it empty for a random seed. |
How the background is kept
Before any voice is recast, the node always splits the recording into a voice track and a music-and-effects track. Preserve background music does not decide whether this split happens. It only decides whether the music track is mixed back under the new voices.
If music bleeds into the new voices, or a voice sounds thin, set Separation quality to Quality — finer separation. Keep Fast when speed matters or the voices already sound clean.
Add a room or an echo to the voices
Voice FX adds one effect to all the new voices together, before the music is mixed back in. The effect sits on the voices only, so with Preserve background music on, the music stays dry under a voice with reverb.
- Reverb spaces: Room (indoor), Bathroom, Car interior, Hall / Lobby, Concert Hall, Church, Cave, Arena / Stadium and Outdoor (open air). Wet / Dry mix from 0 to 100 sets how much room you hear.
- Character: Telephone, Megaphone / PA, Echo / Slap-back and Custom. Echo and Custom use Delay (ms), from 20 to 2,000, and Decay, from 0 to 1, where a higher decay gives more repeats.
The Audio FX page describes each preset.
Re-voice a whole video
Wire any talking video into the Video input. The node then:
- Extracts the sound track from the clip.
- Detects each speaker in the order they first talk.
- Recasts each speaker with the matching voice.
- Puts the new voices back on the original video, and returns the video and the new sound track.
The video needs a sound track. Most text-to-video and image-to-video models make silent video, and a silent clip fails at once. Use a clip with spoken audio, or wire a recording into the Audio input.
Credits
Only recast voices cost credits. Keep-slots and speakers past the end of your list pass through for free.
- A Recast voice is billed by the length of its stem. The stem runs from the start of the clip to that speaker's last line. The rate is 40 credits per minute, counted per second and rounded up to the next credit.
- A Re-speak voice is billed at 30 credits per started 1,000 characters of the text it speaks.
- Minimum. Every voice costs at least 4 credits, which is 6 seconds of speech, and so does every run.
| Recast voices | Credits |
|---|---|
| 1 Recast voice whose last line ends at 26.76 seconds | 18 |
| 1 Recast voice with a 60-second stem | 40 |
| 1 Recast voice with a 61-second stem | 41 |
| 1 Recast voice with a 3-second stem | 4 (the minimum) |
| 2 Recast voices with 60-second stems, and 1 Re-speak voice of 1,500 characters | 40 + 40 + 60 = 140 |
You are charged for the stems that were actually converted, and never more than the amount held when the run started.
Use it as an app
Voice Changer Pro is also a standalone app at voice.nodaro.ai. Paste a link or upload a video, choose a voice for each speaker, and get the converted video back.
Tips
- Order the voices with care. The match is by position: the first voice goes to the first speaker to talk.
- Tune each speaker separately. Each voice has its own settings, so you can steady one speaker and keep another expressive.
- Balance loudness per speaker. Use Volume on each voice: Match source mirrors the original speaker, and Manual sets an exact level.
- Map fewer voices than speakers. Only the listed speakers change; everyone after them keeps their voice.
- Keep the performance. Recast converts speech to speech, so the original acting carries through. Use Re-speak (v3) only when you want a new performance, for example on an audio-only project.
Troubleshooting
One speaker sounds robotic. Lower that voice's Similarity or Stability, or choose a voice whose timbre is closer to the original speaker.
Music bleeds into the voices. Set Separation quality to Quality — finer separation.
The wrong speaker got a voice. The speakers are numbered by who talks first. Check the order with Transcribe and reorder the voices.
A self-hosted run is refused with nodaro_connection_required. The install has no Nodaro Cloud connection. Connect it with OAuth or an API key, as described in Cloud connect.
From the API
POST /v1/voice-changer-pro runs the node from code, on Nodaro Cloud. orderedVoices maps speaker N to entry N. An entry is a voice id, an object with per-voice settings, or null for a keep-slot. The settings object takes voiceId, engine ("sts" or "v3"), stability, similarityBoost, style, useSpeakerBoost, seed, volumeMode and volume. With engine: "v3", stability accepts only 0, 0.5 or 1, and similarityBoost, style and useSpeakerBoost are ignored. The result's sourceHasVideo tells you whether the run used video mode or audio mode.
On Nodaro Cloud, you can also run the same engine in three steps. You can then check the speakers before you pay for a recast, and mix the tracks before you render:
- Analyze.
POST /v1/voice-changer-pro/analyzeseparates the voices from the music and detects the speakers. It returns each speaker's time segments, word count and a transcript snippet, plus the detected language. It costs 10 credits. - Recast to stems. Call
POST /v1/voice-changer-prowithoutput: "stems"and the analysis. The speakers are not detected again, so you can recast with other voices as often as you like. The result is one dry track per voice. - Export.
POST /v1/voice-changer-pro/exportrenders the final video from your mix. Send up to 16 tracks, each with againfrom 0 to 200,muted, and akindofvoiceorbackground, plus an optionalvoiceFx. At least one track must be unmuted. The video stream is copied, never re-encoded. Adjusting the mix is free; the export costs 1 credit.
The SDK methods are client.voices.analyze, client.voices.recast and client.voices.exportMix. The CLI commands are nodaro voice analyze, nodaro voice recast and nodaro voice export. The MCP tools are voice_changer_pro, voice_changer_pro_analyze and voice_changer_pro_export. See Voice and media.
Frequently asked questions
Related
Voice Changer
Transcribe
Dubbing
Audio FX
Connect to Nodaro Cloud
Last updated on
Voice Changer
Replace the voice in a recording or a talking video with another voice, and keep the original emotion, pacing and timing. Speech-to-speech with ElevenLabs.
Voice Design
Create a new voice from a written description, with control over model, loudness, guidance and seed. Get an audio preview and a voice ID you can keep and reuse.