Nodaro Docs
DocumentationNode ReferenceModelsAI Agents (MCP)DevelopersSelf-hostingResearch
Audio

Voice Changer Pro

Detect each speaker in a recording or talking video and give each one a new voice, with the original emotion and timing kept. Priced per minute of speech.

Available on Nodaro Cloud

The Voice Changer Pro node gives every speaker in a conversation a new voice of their own. It detects each speaker in a recording or a talking video, recasts speaker 1 with your first voice, speaker 2 with your second voice, and so on. Each speaker keeps their original emotion, pacing and timing, so a whole dialogue scene, podcast or interview is re-cast in one run.

Voice Changer Pro runs on Nodaro Cloud. On a self-hosted install, the node shows a NODARO mark and runs through your Nodaro Cloud connection, billed to the connected nodaro.ai account.

When to use it

  • Re-cast a dialogue video with several characters in one step.
  • Replace every voice in a podcast or an interview with anonymous voices.
  • Cast a scene with specific character voices, and keep the original performance.
  • Turn a rough recording with several speakers into a polished voiceover.
  • Recast the guests of a panel while the host keeps their own voice.

For a recording with one speaker, Voice Changer is enough.

Quick start

Add the node

Press Tab on the canvas and choose Audio › Voices › Voice Changer Pro.

Wire the recording

Wire an audio node into the Audio input, or a talking video into the Video input.

Add one voice per speaker

Open the settings panel. Use Add a voice once for each speaker, in the order the speakers first talk. For a speaker who should keep their voice, click Keep original — don't recast this speaker instead.

Run it

Click Run on the node. For a video, the re-voiced video appears on the Video output and the new sound track on the Audio output.

videovideoUpload VideoInterview, 3 speakersVoice Changer ProRachel · Keep original · AriaAdd Captions
An interview with three speakers: the host keeps their voice, the two guests are recast, and the result gets captions.

Inputs and outputs

InputAcceptsWhat it does
AudioAudio nodes, such as Upload Audio and Text to DialogueThe recording whose speakers are recast. The node returns audio.
VideoVideo nodes, such as Upload Video and Generate VideoThe talking video to re-voice. The node returns the video with the new voices.

When both inputs are wired, the video wins and the audio input is ignored.

OutputWhat it carries
AudioThe re-voiced audio. Always produced. For a video, this is the new sound track.
VideoThe re-voiced video. Only produced when a video is wired in.

The file decides the mode, not the input. Before any work starts, the node checks what the file really contains:

  • An audio-only file on the Video input, such as a voice recording saved as .mp4 or .m4a, runs as audio. You get the re-voiced audio, the Video output stays empty, and nothing fails.
  • A video on the Audio input has its sound extracted first and is treated as audio.

How speakers get their voices

The node finds each speaker in the recording and numbers them in the order they first talk. Your list of voices is matched to that order.

Your listSpeaker 1Speaker 2Speaker 3Speaker 4
Rachel, Keep original, AriaRachelOriginal voiceAriaOriginal voice
  • Up to 8 voices. The list holds 1 to 8 entries, and at least one entry must be a real voice.
  • Keep-slots. A Keep original entry leaves that speaker's voice unchanged, while later speakers are still recast.
  • Shorter lists. Speakers past the end of the list keep their original voices.
  • Edit the list. Move a voice with the up and down arrows, turn it into a keep-slot with Keep, or remove it.

If you are not sure who talks first, run the recording through Transcribe with Speaker Diarization on, and read the speaker labels.

Settings

SettingWhat it does
Add a voiceOpens the voice browser and adds the chosen voice to the end of the list. Premade, Voice Library and your own voices all work.
Keep original — don't recast this speakerAdds a keep-slot to the end of the list.
ModelMultilingual v2 (the default) covers 29 languages and is ElevenLabs' recommended model, also for English. English v2 is for English only.
Separation qualityFast — preserves more of the voice (the default) or Quality — finer separation.
Preserve background musicOn (the default) mixes the separated music and sound effects back under the new voices. Off drops them for a voice-only result.
Music volumeThe level of the preserved music: Match source (the default) keeps its original level, Normalize levels its loudness, and Manual sets Music level from 0 to 200%. Only when Preserve background music is on.
Voice FXNone (the default), or a reverb or echo on the new voices. See Voice FX.
Remove background noiseOff by default. Denoises the result. The voices are already isolated automatically, so this option may be unnecessary.

Settings for each voice

Open Voice settings under a voice to tune that speaker alone. Leave a setting untouched to use the model's default.

SettingWhat it does
Conversion EngineRecast (the default) converts the speech and keeps the original delivery. Re-speak (v3) generates a new performance from the transcript with ElevenLabs v3, and understands audio tags. The original delivery is replaced, and on a video the lips no longer match.
StabilityFrom 0 to 1 on Recast. On Re-speak: Most Variable (0), Balanced (0.5) or Most Stable (1.0). Higher is steadier, lower is more expressive.
SimilarityFrom 0 to 1. How closely the result follows the target voice's timbre. Recast only.
Style ExaggerationFrom 0 to 1. Above 0, it amplifies the delivery, with more latency and less stability. Recast only.
Speaker BoostSharpens the likeness to the target voice, with slightly more latency. Recast only.
VolumeMatch source (the default) matches the original speaker's loudness. Normalize levels the loudness. Manual sets a level from 0 to 200%.
SeedA whole number from 0 to 4,294,967,295. The same source, settings and seed recast this speaker identically. Leave it empty for a random seed.

How the background is kept

Before any voice is recast, the node always splits the recording into a voice track and a music-and-effects track. Preserve background music does not decide whether this split happens. It only decides whether the music track is mixed back under the new voices.

If music bleeds into the new voices, or a voice sounds thin, set Separation quality to Quality — finer separation. Keep Fast when speed matters or the voices already sound clean.

Add a room or an echo to the voices

Voice FX adds one effect to all the new voices together, before the music is mixed back in. The effect sits on the voices only, so with Preserve background music on, the music stays dry under a voice with reverb.

  • Reverb spaces: Room (indoor), Bathroom, Car interior, Hall / Lobby, Concert Hall, Church, Cave, Arena / Stadium and Outdoor (open air). Wet / Dry mix from 0 to 100 sets how much room you hear.
  • Character: Telephone, Megaphone / PA, Echo / Slap-back and Custom. Echo and Custom use Delay (ms), from 20 to 2,000, and Decay, from 0 to 1, where a higher decay gives more repeats.

The Audio FX page describes each preset.

Re-voice a whole video

Wire any talking video into the Video input. The node then:

  1. Extracts the sound track from the clip.
  2. Detects each speaker in the order they first talk.
  3. Recasts each speaker with the matching voice.
  4. Puts the new voices back on the original video, and returns the video and the new sound track.

The video needs a sound track. Most text-to-video and image-to-video models make silent video, and a silent clip fails at once. Use a clip with spoken audio, or wire a recording into the Audio input.

Credits

Only recast voices cost credits. Keep-slots and speakers past the end of your list pass through for free.

  • A Recast voice is billed by the length of its stem. The stem runs from the start of the clip to that speaker's last line. The rate is 40 credits per minute, counted per second and rounded up to the next credit.
  • A Re-speak voice is billed at 30 credits per started 1,000 characters of the text it speaks.
  • Minimum. Every voice costs at least 4 credits, which is 6 seconds of speech, and so does every run.
Recast voicesCredits
1 Recast voice whose last line ends at 26.76 seconds18
1 Recast voice with a 60-second stem40
1 Recast voice with a 61-second stem41
1 Recast voice with a 3-second stem4 (the minimum)
2 Recast voices with 60-second stems, and 1 Re-speak voice of 1,500 characters40 + 40 + 60 = 140

You are charged for the stems that were actually converted, and never more than the amount held when the run started.

Use it as an app

Voice Changer Pro is also a standalone app at voice.nodaro.ai. Paste a link or upload a video, choose a voice for each speaker, and get the converted video back.

Tips

  • Order the voices with care. The match is by position: the first voice goes to the first speaker to talk.
  • Tune each speaker separately. Each voice has its own settings, so you can steady one speaker and keep another expressive.
  • Balance loudness per speaker. Use Volume on each voice: Match source mirrors the original speaker, and Manual sets an exact level.
  • Map fewer voices than speakers. Only the listed speakers change; everyone after them keeps their voice.
  • Keep the performance. Recast converts speech to speech, so the original acting carries through. Use Re-speak (v3) only when you want a new performance, for example on an audio-only project.

Troubleshooting

One speaker sounds robotic. Lower that voice's Similarity or Stability, or choose a voice whose timbre is closer to the original speaker.

Music bleeds into the voices. Set Separation quality to Quality — finer separation.

The wrong speaker got a voice. The speakers are numbered by who talks first. Check the order with Transcribe and reorder the voices.

A self-hosted run is refused with nodaro_connection_required. The install has no Nodaro Cloud connection. Connect it with OAuth or an API key, as described in Cloud connect.

From the API

POST /v1/voice-changer-pro runs the node from code, on Nodaro Cloud. orderedVoices maps speaker N to entry N. An entry is a voice id, an object with per-voice settings, or null for a keep-slot. The settings object takes voiceId, engine ("sts" or "v3"), stability, similarityBoost, style, useSpeakerBoost, seed, volumeMode and volume. With engine: "v3", stability accepts only 0, 0.5 or 1, and similarityBoost, style and useSpeakerBoost are ignored. The result's sourceHasVideo tells you whether the run used video mode or audio mode.

On Nodaro Cloud, you can also run the same engine in three steps. You can then check the speakers before you pay for a recast, and mix the tracks before you render:

  1. Analyze. POST /v1/voice-changer-pro/analyze separates the voices from the music and detects the speakers. It returns each speaker's time segments, word count and a transcript snippet, plus the detected language. It costs 10 credits.
  2. Recast to stems. Call POST /v1/voice-changer-pro with output: "stems" and the analysis. The speakers are not detected again, so you can recast with other voices as often as you like. The result is one dry track per voice.
  3. Export. POST /v1/voice-changer-pro/export renders the final video from your mix. Send up to 16 tracks, each with a gain from 0 to 200, muted, and a kind of voice or background, plus an optional voiceFx. At least one track must be unmuted. The video stream is copied, never re-encoded. Adjusting the mix is free; the export costs 1 credit.

The SDK methods are client.voices.analyze, client.voices.recast and client.voices.exportMix. The CLI commands are nodaro voice analyze, nodaro voice recast and nodaro voice export. The MCP tools are voice_changer_pro, voice_changer_pro_analyze and voice_changer_pro_export. See Voice and media.

Frequently asked questions

Last updated on

On this page