Nodaro Docs
DocumentationNode ReferenceModelsAI Agents (MCP)DevelopersSelf-hostingResearch
Audio models

Audio models

Every audio model you can run in Nodaro, with what each one does and what it costs in credits.

Nodaro runs 20 audio models. Each row below links to a page with the model's settings, credit prices and prompt tips.

ElevenLabs

ModelMakerModesCreditsDetails
ElevenLabs v3ElevenLabsText to speech30Latest ElevenLabs TTS — supports [audio tags] for emotion / pacing. Direct API.
ElevenLabs Turbo v2.5ElevenLabsText to speech15Fast, cheap ElevenLabs TTS via the direct ElevenLabs API. Good for narration.
ElevenLabs Multilingual v2ElevenLabsText to speech30Multi-language ElevenLabs TTS via the direct ElevenLabs API.
ElevenLabs Dialogue v3ElevenLabsMulti-speaker dialogue25Multi-speaker dialogue via the direct ElevenLabs API — give it a script, it voices each role (any voice: premade, library, or cloned).
ElevenLabs Voice DesignElevenLabsVoice design50Design a synthetic voice from a description (no reference clip needed).
ElevenLabs Voice ChangerElevenLabsVoice changer40Speech-to-speech: convert one voice to another while preserving prosody.
ElevenLabs STTElevenLabsSpeech to text22Speech-to-text with WORD-level timestamps (always on), speaker diarization and audio-event tags. The engine to use when the transcript feeds captions.
ElevenLabs Voice IsolationElevenLabsVoice isolation74Strip background noise / music from a vocal track.
ElevenLabs DubbingElevenLabsDubbing40Translate + dub audio or a whole video into a new language — video in, dubbed video out. Async.
ElevenLabs Dubbing v2ElevenLabsDubbing1100Translate + dub audio or a whole video into a new language — video in, dubbed video out. Async.
ElevenLabs Forced AlignmentElevenLabsForced alignment30Align an existing transcript to audio with word-level timestamps.
ElevenLabs Sound EffectsElevenLabsSound effects3Generate short sound effects from a text prompt.

OpenAI

ModelMakerModesCreditsDetails
Incredibly Fast WhisperOpenAISpeech to text40Fast Whisper speech-to-text. Returns WORD-level timestamps when asked, so its transcript can feed captions.
WhisperOpenAISpeech to text40Whisper speech-to-text — PHRASE-level segments only, NO word timestamps. Fine for a transcript or a static subtitle; not for word-timed (kinetic) captions.

Suno

ModelMakerModesCreditsDetails
Suno v4SunoMusic30Suno v4 music generation — full songs with vocals, multiple genres.
Suno v5SunoMusic30Suno v5 — better vocal quality than v4, more genres. Same price.
Suno v5.5SunoMusic30Suno v5.5 — improved audio quality and expressiveness over v5.
Suno V6SunoMusic30Suno V6 — greater musical expression with more natural vocals and richer details. The flagship and the default.
Suno V6 WildSunoMusic30Suno V6 Wild — pushes creative boundaries for bolder, more distinctive musical expression; more varied, less predictable results.
Suno V6 MiniSunoMusic30Suno V6 Mini — lightweight and fast, balancing quality and speed for effortless creation.

Last updated on