# ボイスとオーディオ

> Nodaro SDK を使って、TypeScript からボイスの参照、変更、差し替え、デザイン、動画の吹き替え、オーディオの分離、ミックス、文字起こしを行います。

Source: https://nodaro.ai/ja/docs/developers/sdk/voices-and-audio

**`client.voices`** は ElevenLabs のボイスを扱います。ボイスの一覧表示と検索、録音の声の変更、話者ごとの新しいボイスの割り当て、新しいボイスのデザイン、オーディオや動画の別の言語への吹き替えができます。**`client.audio`** は、オーディオの構成要素をそれぞれ単独で提供します。分離、声の抽出、エフェクト、ミックス、音量、結合、文字起こしです。どの生成メソッドも `jobId` を返すので、[`client.jobs.getStatus()`](https://nodaro.ai/docs/developers/sdk/jobs-and-executions) でポーリングしてください。これらのメソッドは、[音声とメディアの REST API](https://nodaro.ai/docs/developers/api/voice-and-media) を呼び出します。

## メソッド
| メソッド | 内容 |
| --- | --- |
| [`voices.list()`](#voiceslist) | 標準ボイスを一覧表示します |
| [`voices.searchLibrary(params?)`](#voicessearchlibraryparams) | コミュニティのボイスライブラリを検索します |
| [`voices.listClones()`](#voiceslistclones) | 既存のボイスクローンを一覧表示します |
| [`voices.deleteClone(id)`](#voicesdeletecloneid) | ボイスクローンの 1 つを削除します |
| [`voices.createClone()` と `createCloneFromFile()`](#voicescreateclone-and-createclonefromfile) | 廃止されました |
| [`voices.change(input)`](#voiceschangeinput) | 録音や動画の声を差し替えます |
| [`voices.recast(input)`](#voicesrecastinput) | 話者ごとに異なるボイスを割り当てます |
| [`voices.analyze(input)`](#voicesanalyzeinput) | 差し替えの前に話者を検出します |
| [`voices.exportMix(input)`](#voicesexportmixinput) | ミックスしたボイスのステムから動画をレンダリングします |
| [`voices.design(input)`](#voicesdesigninput) | 説明から新しいボイスを作成します |
| [`voices.remix(input)`](#voicesremixinput) | 説明したボイスでテキストを話します |
| [`voices.dub(input)`](#voicesdubinput) | オーディオや動画を別の言語に吹き替えます |
| [`voices.textToDialogue(input)`](#voicestexttodialogueinput) | 複数話者の脚本を 1 つのオーディオファイルとして音声にします |
| [`audio.separate(input)`](#audioseparateinput) | トラックをステムに分割します |
| [`audio.isolate(input)`](#audioisolateinput) | 主な声を残し、ノイズを取り除きます |
| [`audio.applyFx(input)`](#audioapplyfxinput) | リバーブ、エコー、電話、メガホンを加えます |
| [`audio.mix(input)`](#audiomixinput) | 複数のトラックを重ねて 1 つにします |
| [`audio.adjustVolume(input)`](#audioadjustvolumeinput) | レベルの変更、ノーマライズ、フェードを行います |
| [`audio.combine(input)`](#audiocombineinput) | オーディオのセグメントを順番につなげます |
| [`audio.transcribe(input)`](#audiotranscribeinput) | 単語のタイミング付きで、音声をテキストにします |

## client.voices：ボイス
### voices.list()
ElevenLabs の標準ボイスを一覧表示します（`GET /v1/voices`）。サーバーに ElevenLabs のキーが設定されていない場合は、代わりに厳選されたセットを返します。

```ts
list(): Promise<Voice[]>
```

```ts
const voices = await client.voices.list()
```

### voices.searchLibrary(params?)
コミュニティのボイスライブラリを検索します（`GET /v1/voices/library`）。どのパラメーターも任意で、空の値は取り除かれるため、サーバーのデフォルトが適用されます。応答の `hasMore` で、次のページがあるかどうかがわかります。

```ts
searchLibrary(params?: VoiceLibraryParams): Promise<{ voices: Voice[]; hasMore: boolean }>
```

<TypeTable
type={{
search: { type: 'string', description: "検索する語句です。" },
gender: { type: 'string', description: "性別で絞り込みます。" },
age: { type: 'string', description: "年齢で絞り込みます。" },
accent: { type: 'string', description: "アクセントで絞り込みます。" },
language: { type: 'string', description: "言語コードで絞り込みます。en などです。" },
category: { type: 'string', description: "カテゴリーで絞り込みます。" },
use_cases: { type: 'string', description: "用途で絞り込みます。" },
descriptives: { type: 'string', description: "特徴を表すタグで絞り込みます。" },
featured: { type: 'boolean', description: "おすすめのボイスのみです。" },
sort: { type: 'string', description: "並び順です。trending などです。" },
page: { type: 'number', default: '0', description: "ページ番号で、0 から数えます。" },
page_size: { type: 'number', default: '30', description: "1 ページあたりのボイスの数で、1〜100 です。" },
}}
/>

```ts
const { voices, hasMore } = await client.voices.searchLibrary({ search: "deep", language: "en" })
const voice = voices[0]

await client.nodes.run("text-to-speech", {
text: "Hello!",
voice: voice.voice_id,
voiceType: "library",
...(voice.recommendedProvider ? { provider: voice.recommendedProvider } : {}),
})
```

各ボイスには、そのボイスが検証済みの音声モデルについて、2 つのヒントが付くことがあります。

- **`recommendedProvider`** は、そのボイスに最適なモデルです。モデルピッカーのないアプリは、これを `provider` として[**テキストから音声**（Text to Speech）](https://nodaro.ai/docs/nodes/audio/text-to-speech)に送ってください。そうすると、ボイスがライブラリのプレビューと同じように聞こえます。
- **`verifiedProviders`** は、そのボイスが検証済みのすべてのモデルを一覧にします。モデルピッカーのあるアプリは、ユーザーの選んだモデルがこの一覧にない場合にだけ、それを置き換えてください。

### voices.listClones()
自分のボイスクローンを一覧表示します（`GET /v1/voice-clones`）。クローン機能が廃止される前に作ったクローンは、ボイスを指定できるところならどこでも、ボイス ID として引き続き使えます。

```ts
listClones(): Promise<VoiceClone[]>
```

```ts
const clones = await client.voices.listClones()
```

### voices.deleteClone(id)
自分のボイスクローンの 1 つを削除します（`DELETE /v1/voice-clones/:id`）。

```ts
deleteClone(id: string): Promise<void>
```

<TypeTable
type={{
id: { type: 'string', required: true, description: "クローンの ID です。" },
}}
/>

```ts
await client.voices.deleteClone(cloneId)
```

### voices.createClone() と createCloneFromFile()
ボイスクローンの作成は、Nodaro ではもう提供されていません。2026 年 9 月に廃止されました。両方のメソッドはクライアントに残っていますが、非推奨として示されます。`code` が `voice_cloning_retired` の `NodaroError` を、ステータス 410 でスローします。新しいカスタムボイスには、代わりに [`design()`](#voicesdesigninput) を使ってください。

## client.voices：ボイスチェンジャー
### voices.change(input)
録音や、話している動画全体の声を、別の声に置き換えます（`POST /v1/voice-changer`）。[**ボイスチェンジャー**（Voice Changer）](https://nodaro.ai/docs/nodes/audio/voice-changer)ノードと同じ処理です。`videoUrl` を指定すると、サーバーは動画から音声を取り出し、声を変換して、新しい声を元の映像に載せ直します。

```ts
change(input: {
voiceId: string
audioUrl?: string
videoUrl?: string
model?: string
stability?: number
similarityBoost?: number
style?: number
useSpeakerBoost?: boolean
seed?: number
removeBackgroundNoise?: boolean
}): Promise<{ jobId: string }>
```

<TypeTable
type={{
voiceId: { type: 'string', required: true, description: "新しいボイスです。標準ボイスの名前か、ElevenLabs のボイス ID です。" },
audioUrl: { type: 'string', description: "変換する録音です。audioUrl か videoUrl を指定します。両方を送った場合は、動画が優先されます。" },
videoUrl: { type: 'string', description: "声を変換する動画です。" },
model: { type: 'string', description: "声質変換のモデルを上書きします。" },
stability: { type: 'number', description: "ボイスの安定性で、0〜1 です。" },
similarityBoost: { type: 'number', description: "変換先のボイスにどれだけ近づけるかです。0〜1 です。" },
style: { type: 'number', default: '0', description: "スタイルの誇張で、0〜1 です。0 より大きくすると話し方を強調しますが、速度と安定性が少し落ちます。" },
useSpeakerBoost: { type: 'boolean', description: "変換先のボイスへの類似性を高めます。わずかに遅くなります。" },
seed: { type: 'number', description: "結果を再現できるようにする整数です。" },
removeBackgroundNoise: { type: 'boolean', description: "true は声だけのクリーンな結果を返します。false は音楽と効果音を新しい声の下に残します。" },
}}
/>

```ts
const { jobId } = await client.voices.change({
videoUrl: "https://example.com/talking.mp4",
voiceId: "Aria",
})
// the finished job's output_data has videoUrl and audioUrl
```

### voices.recast(input)
録音の中で検出された各話者に、それぞれ異なるボイスを割り当てます（`POST /v1/voice-changer-pro`）。[**ボイスチェンジャー Pro**（Voice Changer Pro）](https://nodaro.ai/docs/nodes/audio/voice-changer-pro)ノードと同じ処理です。Nodaro Cloud で実行され、クレジットがかかり、ジョブとして実行されます。

```ts
recast(input: VoiceChangerProInput): Promise<{ jobId: string }>
```

<TypeTable
type={{
orderedVoices: { type: 'Array<VoiceChangerProVoice | null>', required: true, description: "1〜8 件のエントリーで、検出された話者の順に 1 件ずつ対応します。各エントリーは、ボイス ID、話者ごとの設定を持つオブジェクト、またはその話者の声をそのまま維持する null のいずれかです。少なくとも 1 件は null 以外にする必要があります。" },
audioUrl: { type: 'string', description: "録音です。audioUrl か videoUrl を指定します。" },
videoUrl: { type: 'string', description: "話者の声を差し替える動画です。新しい声は、元の映像に載せ直されます。" },
model: { type: 'string', description: "声質変換のモデルを上書きします。" },
preserveBackground: { type: 'boolean', default: 'true', description: "分離した音楽と効果音を、新しい声の下に再びミックスします。false の場合は、声だけを返します。" },
separationQuality: { type: '"fast" | "best"', default: '"fast"', description: "fast は高速で、声をより多く残します。best は声と音楽をより精密に分離します。" },
removeBackgroundNoise: { type: 'boolean', description: "結果からノイズも取り除きます。" },
musicVolumeMode: { type: '"match" | "normalize" | "manual"', default: '"match"', description: "残す背景音のレベルです。元のレベル、ノーマライズ、または musicVolume のいずれかです。" },
musicVolume: { type: 'number', description: "背景音のレベルをパーセントで指定します。0〜200 で、musicVolumeMode が manual のときに使います。" },
voiceFx: { type: '{ preset, wetDryMix?, delayMs?, decay? }', description: "新しいすべての声にかけるエフェクトです。背景音が戻る前に適用されます。" },
output: { type: '"video" | "stems"', default: '"video"', description: "video は完成した結果をレンダリングします。stems は、自分でミックスして exportMix() でレンダリングするための、別々の未ミックスのトラックを返します。" },
analysis: { type: 'VcpAnalysis', description: "話者の検出をスキップするための、以前の analyze() ジョブの結果です。" },
}}
/>

`orderedVoices` の各エントリーは、次のいずれかです。

- **ボイス ID**：`"Rachel"` のような標準ボイスの名前か、ElevenLabs のボイス ID です。
- **`null`**：元の声を維持するスロットです。その話者は自分の声のままですが、それより後の話者は引き続き差し替えられます。差し替える話者にだけ料金がかかるため、維持するスロットは無料です。
- **オブジェクト**：`voiceId` と、その話者の設定です。`stability`、`similarityBoost`、`style` は 0〜1、`useSpeakerBoost`、`seed` は 0〜4,294,967,295 です。`volumeMode` は `"match"`（デフォルト、元の話者のレベル）、`"normalize"`、または `volume` を 0〜200% で指定する `"manual"` のいずれかです。`engine: "v3"` は、録音した音声を変換する代わりに、その話者のセリフを文字起こしから読み上げ直します。`[audio tags]` にも対応します。

話者 0 には `orderedVoices[0]` が、話者 1 には `orderedVoices[1]` が割り当てられ、以降も同様です。最後のエントリーより後の話者は、自分の声のままです。声と音楽は、常に最初に分離されます。`preserveBackground` が決めるのは、音楽を戻すかどうかだけです。

`voiceFx.preset` は、リバーブの空間（`room`、`bathroom`、`car`、`hall`、`concert-hall`、`church`、`cave`、`arena`、`outdoor`）か、`telephone`、`megaphone`、`echo`、`custom` のいずれかです。リバーブのプリセットは、0〜100 の `wetDryMix` を使います。`echo` と `custom` は、20〜2,000 の `delayMs` と、0〜1 の `decay` を使います。

```ts
// Recast speakers 1 and 3, and keep speaker 2's own voice
const { jobId } = await client.voices.recast({
audioUrl: "https://example.com/panel.mp3",
orderedVoices: ["Rachel", null, "Aria"],
})

// Repeatable voices, a hall reverb, and finer separation
const { jobId: tuned } = await client.voices.recast({
audioUrl: "https://example.com/dialogue.mp3",
orderedVoices: [
{ voiceId: "Rachel", seed: 12345, stability: 0.6 },
{ voiceId: "Aria", seed: 67890, volumeMode: "manual", volume: 120 },
],
voiceFx: { preset: "hall", wetDryMix: 35 },
separationQuality: "best",
})
```

動画モードでは、完了したジョブの `output_data` に `videoUrl` と `audioUrl` が入ります。

### voices.analyze(input)
クリップの話者を、**差し替えずに**検出します（`POST /v1/voice-changer-pro/analyze`）。声を音楽から 1 回分離し、誰がいつ話しているかを調べます。Nodaro Cloud で実行され、料金は一律で、ジョブとして実行されます。

```ts
analyze(input: {
audioUrl?: string
videoUrl?: string
separationQuality?: "fast" | "best"
suggestTitle?: boolean
}): Promise<{ jobId: string }>
```

<TypeTable
type={{
audioUrl: { type: 'string', description: "録音です。audioUrl と videoUrl のどちらか一方だけを指定します。" },
videoUrl: { type: 'string', description: "分析する動画です。" },
separationQuality: { type: '"fast" | "best"', description: "分離の品質です。recast() と同じです。" },
suggestTitle: { type: 'boolean', description: "クリップのタイトルも提案します。" },
}}
/>

完了したジョブの `output_data` は `VcpAnalysis` です。分離した `vocalsUrl` と `backgroundUrl`、検出された `speakers`（それぞれに `id`、時間の `segments`、`firstStartSec`、`wordCount`、テキストの `snippet` を含む）、`languageProbability` を伴う `languageCode`、依頼した場合は `suggestedTitle` が含まれます。これを、後で呼び出す `recast()` に `analysis` として渡すと、分離済みのトラックが再利用され、検出のための料金がもう一度かかることはありません。

### voices.exportMix(input)
ミックスしたステムから、最終的な動画をレンダリングします（`POST /v1/voice-changer-pro/export`）。`recast({ output: "stems" })` の後に行う、最後のステップです。Nodaro Cloud で実行され、料金は一律で、ジョブとして実行されます。

```ts
exportMix(input: {
videoUrl: string
tracks: Array<{ url: string; gain: number; muted: boolean; kind?: "voice" | "background" }>
voiceFx?: { preset: AudioFxPreset; wetDryMix?: number; delayMs?: number; decay?: number }
}): Promise<{ jobId: string }>
```

<TypeTable
type={{
videoUrl: { type: 'string', required: true, description: "動画です。その映像は、再エンコードせずにコピーされます。" },
tracks: { type: 'VcpExportTrack[]', required: true, description: "最大 16 本のトラックで、少なくとも 1 本はミュートを解除しておく必要があります。それぞれに url、0〜200 の gain、muted、任意の kind があります。" },
voiceFx: { type: '{ preset, wetDryMix?, delayMs?, decay? }', description: "声のトラックだけにかけるエフェクトです。背景のトラックには、かかりません。" },
}}
/>

対話的なフローでは、1 回分析し、ステムに差し替え、その後ミックスしてレンダリングします。以下は、簡単なポーリング用のヘルパーを使った例です。

```ts

async function outputOf(jobId: string): Promise<any> {
for (;;) {
const { data } = await client.jobs.getStatus(jobId)
if (data.status === "completed") return data.output_data
if (data.status === "failed" || data.status === "cancelled") throw new Error(data.error_message ?? data.status)
await new Promise((resolve) => setTimeout(resolve, 2_000))
}
}

const { jobId: analyzeJob } = await client.voices.analyze({ videoUrl })
const analysis = (await outputOf(analyzeJob)) as VcpAnalysis

const { jobId: recastJob } = await client.voices.recast({
videoUrl,
orderedVoices: ["Rachel", null, "Aria"],
output: "stems",
analysis,
})
const stems = await outputOf(recastJob)

const { jobId: exportJob } = await client.voices.exportMix({
videoUrl,
tracks: [
{ url: stems.tracks[0].url, gain: 100, muted: false },
{ url: stems.tracks[1].url, gain: 90, muted: false },
{ url: stems.backgroundUrl, gain: 70, muted: false, kind: "background" },
],
voiceFx: { preset: "hall", wetDryMix: 25 },
})
const { videoUrl: finalVideo } = await outputOf(exportJob)
```

ゲイン、ミュート、エフェクトは、動画をレンダリングするときに適用され、映像はそのままコピーされます。そのため、エクスポートの結果はプレビューと一致し、エクスポートするまでは、好きなだけミックスを変更できます。すべてのトラックをミュートしたミックスは、400 で拒否されます。

## client.voices：作成と翻訳
### voices.design(input)
文章による説明から、新しい合成ボイスを作成します（`POST /v1/voice-design`）。[**ボイスデザイン**（Voice Design）](https://nodaro.ai/docs/nodes/audio/voice-design)ノードと同じ処理です。完了したジョブには、オーディオのプレビューと、新しいボイスの ID が含まれ、ボイスを指定できるところならどこでも使えます。

```ts
design(input: {
text: string
voiceDescription: string
model?: string
loudness?: number
guidanceScale?: number
seed?: number
quality?: number
shouldEnhance?: boolean
userPrompt?: string
}): Promise<{ jobId: string }>
```

<TypeTable
type={{
text: { type: 'string', required: true, description: "プレビューが話す文で、100〜1,000 文字です。" },
voiceDescription: { type: 'string', required: true, description: "ほしいボイスです。たとえば、年配のナレーターの温かく低い声などです。" },
model: { type: 'string', description: "ボイスデザインのモデルを上書きします。" },
loudness: { type: 'number', description: "ラウドネスで、-1〜1 です。" },
guidanceScale: { type: 'number', description: "説明にどれだけ厳密に従うかです。0〜100 です。" },
seed: { type: 'number', description: "結果を再現できるようにするシードです。" },
quality: { type: 'number', description: "品質の設定です。" },
shouldEnhance: { type: 'boolean', description: "モデルに説明を拡張させます。" },
userPrompt: { type: 'string', description: "もとのリクエストで、ジョブとともに保存されます。" },
}}
/>

```ts
const { jobId } = await client.voices.design({
text: "Welcome back. Tonight we follow the river north, into the mountains where the story began.",
voiceDescription: "A calm, deep voice of an older male narrator with a slight British accent",
})
```

### voices.remix(input)
わかりやすい言葉で説明したボイスでテキストを話します。ボイスを作成することはありません（`POST /v1/voice-remix`）。[**ボイスリミックス**（Voice Remix）](https://nodaro.ai/docs/nodes/audio/voice-remix)ノードと同じ処理です。

```ts
remix(input: { text: string; voiceDescription: string; userPrompt?: string }): Promise<{ jobId: string }>
```

<TypeTable
type={{
text: { type: 'string', required: true, description: "話すテキストで、1〜5,000 文字です。" },
voiceDescription: { type: 'string', required: true, description: "使うボイスを、言葉で説明したものです。" },
userPrompt: { type: 'string', description: "もとのリクエストで、ジョブとともに保存されます。" },
}}
/>

```ts
const { jobId } = await client.voices.remix({
text: "Your order is on its way.",
voiceDescription: "A cheerful young woman, fast and upbeat",
})
```

### voices.dub(input)
オーディオ、または動画全体を、各話者の声を保ったまま別の言語に吹き替えます（`POST /v1/dubbing`）。[**吹き替え**（Dubbing）](https://nodaro.ai/docs/nodes/audio/dubbing)ノードと同じ処理です。動画の吹き替えでは、吹き替えたクリップである `output_data.videoUrl` と、吹き替えたトラックのみの `output_data.audioUrl` が返ります。

```ts
dub(input: DubbingInput): Promise<{ jobId: string }>
```

<TypeTable
type={{
targetLanguage: { type: 'string', required: true, description: "吹き替え先の言語で、es や pt-BR のようなコードです。" },
audioUrl: { type: 'string', description: "オーディオのソースです。audioUrl、videoUrl、sourceUrl のうち、ちょうど 1 つを指定します。" },
videoUrl: { type: 'string', description: "動画のソースです。結果は吹き替えた動画になります。" },
sourceUrl: { type: 'string', description: "公開されている YouTube、TikTok、または直接のリンクで、代わりに取得されます。" },
sourceLanguage: { type: 'string', description: "話されている言語です。省略すると自動検出されます。" },
numSpeakers: { type: 'number', default: '0', description: "話者の数で、1〜20、または自動検出する場合は 0 です。人数がわかっていると、分離の精度が上がります。" },
disableVoiceCloning: { type: 'boolean', description: "各話者本人の声の代わりに、標準ボイスを使います。" },
dropBackgroundAudio: { type: 'boolean', description: "音楽と効果音を取り除きます。" },
startTime: { type: 'number', description: "この時点から吹き替えます。単位は秒です。" },
endTime: { type: 'number', description: "この時点まで吹き替えます。単位は秒です。" },
highestResolution: { type: 'boolean', description: "動画の吹き替えで、ソースの解像度を維持します。" },
useProfanityFilter: { type: 'boolean', description: "不適切な表現を除去します。" },
targetAccent: { type: 'string', description: "吹き替えのアクセントです。試験的な機能です。" },
watermark: { type: 'boolean', description: "動画の吹き替えに、モデルの開発元の透かしを追加します。" },
}}
/>

```ts
const { jobId } = await client.voices.dub({
videoUrl: "https://example.com/interview.mp4",
targetLanguage: "es",
numSpeakers: 2,
})
```

料金は、吹き替える区間の分数によって決まり、最低 1 分として計算されます。区間は最長 30 分なので、それより長いソースには `startTime` と `endTime` を使ってください。

### voices.textToDialogue(input)
複数の話者による脚本を、1 つのオーディオファイルとして音声にします（`POST /v1/text-to-dialogue`）。[**テキストから会話**（Text to Dialogue）](https://nodaro.ai/docs/nodes/audio/text-to-dialogue)ノードが [ElevenLabs Dialogue v3](https://nodaro.ai/docs/models/audio/elevenlabs-dialogue-v3) で行うのと同じ処理です。

```ts
textToDialogue(input: {
dialogue: Array<{ text: string; voice: string }>
stability?: 0 | 0.5 | 1
languageCode?: string
seed?: number
applyTextNormalization?: "auto" | "on" | "off"
}): Promise<{ jobId: string }>
```

<TypeTable
type={{
dialogue: { type: 'Array<{ text: string; voice: string }>', required: true, description: "話す順に並んだセリフです。voice は標準ボイスの名前か、ElevenLabs のボイス ID です。テキストには [laughs] のようなオーディオタグを含められます。" },
stability: { type: '0 | 0.5 | 1', description: "話し方の安定性です。" },
languageCode: { type: 'string', description: "ISO 639-1 の言語のヒントです。省略すると自動検出されます。" },
seed: { type: 'number', description: "結果を再現できるようにする場合は 0〜4,294,967,295 です。ランダムな結果にするには省略します。" },
applyTextNormalization: { type: '"auto" | "on" | "off"', description: "数字や略語を読み方どおりに展開するかどうかです。" },
}}
/>

```ts
const { jobId } = await client.voices.textToDialogue({
dialogue: [
{ text: "Did you hear that?", voice: "Rachel" },
{ text: "[whispers] Stay behind me.", voice: "Callum" },
],
})
```

脚本は最大 5,000 文字まで使え、2,000 文字未満だと最も高品質になります。使えるボイスは最大 10 種類です。ボイスライブラリのボイス、デザインしたボイス、既存のクローンはどれも使えて、自由に組み合わせられます。完了したジョブの `output_data.audioUrl` が、そのファイルです。

## client.audio
ボイスチェンジャー Pro が内部で使っているオーディオの構成要素で、1 つずつ個別に使えます。どのメソッドも `jobId` を返します。

### audio.separate(input)
トラックをステムに分割します（`POST /v1/audio-separation`）。[**音源分離**（Audio Separation）](https://nodaro.ai/docs/nodes/audio/audio-separation)ノードと同じ処理です。

```ts
separate(input: { audioUrl: string; mode?: "vocal_instrumental" | "stems"; quality?: "auto" | "fast" | "best" }): Promise<{ jobId: string }>
```

<TypeTable
type={{
audioUrl: { type: 'string', required: true, description: "分割するトラックです。" },
mode: { type: '"vocal_instrumental" | "stems"', default: '"vocal_instrumental"', description: "vocal_instrumental は、声を音楽と効果音から分離します。stems は、ドラム、ベースなどのパートを返します。" },
quality: { type: '"auto" | "fast" | "best"', description: "分離の品質です。" },
}}
/>

```ts
const { jobId } = await client.audio.separate({ audioUrl: songUrl })
// output: vocalUrl and instrumentalUrl, or one URL per stem in stems mode
```

### audio.isolate(input)
主な声を残し、背景ノイズを取り除きます（`POST /v1/audio-isolation`）。[**音声抽出**（Voice Extractor）](https://nodaro.ai/docs/nodes/audio/voice-extractor)ノードと同じ処理です。

```ts
isolate(input: { audioUrl: string }): Promise<{ jobId: string }>
```

<TypeTable
type={{
audioUrl: { type: 'string', required: true, description: "クリーンにする録音です。" },
}}
/>

```ts
const { jobId } = await client.audio.isolate({ audioUrl: interviewUrl })
```

### audio.applyFx(input)
リバーブ、エコー、電話、メガホンのエフェクトを加えます（`POST /v1/audio-fx`）。[**オーディオ FX**（Audio FX）](https://nodaro.ai/docs/nodes/audio/audio-fx)ノードと同じ処理です。プリセットは、ボイスチェンジャーの `voiceFx` と同じです。

```ts
applyFx(input: {
audioUrl: string
preset?: AudioFxPreset
mix?: number
delayMs?: number
decay?: number
eqLow?: number
eqHigh?: number
}): Promise<{ jobId: string }>
```

<TypeTable
type={{
audioUrl: { type: 'string', required: true, description: "処理するトラックです。" },
preset: { type: 'AudioFxPreset', description: "room、hall、church のようなリバーブの空間か、telephone、megaphone、echo、custom のいずれかです。" },
mix: { type: 'number', description: "リバーブのウェット／ドライ比で、0〜100 です。" },
delayMs: { type: 'number', description: "エコーのディレイです。echo と custom で使います。" },
decay: { type: 'number', description: "エコーのディケイです。echo と custom で使います。" },
eqLow: { type: 'number', description: "低域のカットまたはブーストで、単位は dB です。telephone と megaphone で使います。" },
eqHigh: { type: 'number', description: "高域のカットまたはブーストで、単位は dB です。telephone と megaphone で使います。" },
}}
/>

```ts
const { jobId } = await client.audio.applyFx({ audioUrl: lineUrl, preset: "telephone" })
```

### audio.mix(input)
複数のトラックを重ねて 1 つにします（`POST /v1/mix-audio`）。[**オーディオをミックス**（Mix Audio）](https://nodaro.ai/docs/nodes/audio/mix-audio)ノードと同じ処理です。

```ts
mix(input: { audioUrls: string[]; trackVolumes?: number[] }): Promise<{ jobId: string }>
```

<TypeTable
type={{
audioUrls: { type: 'string[]', required: true, description: "2〜20 個のトラックで、同時に再生されます。" },
trackVolumes: { type: 'number[]', description: "トラックごとのレベルをパーセントで指定します。0〜200 で、トラックと同じ順番です。" },
}}
/>

```ts
const { jobId } = await client.audio.mix({ audioUrls: [voiceUrl, musicUrl], trackVolumes: [100, 35] })
```

### audio.adjustVolume(input)
オーディオファイル、または動画の音の音量を変更します（`POST /v1/adjust-volume`）。[**音量調整**（Adjust Volume）](https://nodaro.ai/docs/nodes/audio/adjust-volume)ノードと同じ処理です。

```ts
adjustVolume(input: {
audioUrl?: string
videoUrl?: string
volume?: number
normalize?: boolean
fadeIn?: number
fadeOut?: number
}): Promise<{ jobId: string }>
```

<TypeTable
type={{
audioUrl: { type: 'string', description: "オーディオファイルです。audioUrl か videoUrl を指定します。" },
videoUrl: { type: 'string', description: "音を変更する動画です。" },
volume: { type: 'number', default: '100', description: "レベルをパーセントで指定します。" },
normalize: { type: 'boolean', description: "ラウドネスを標準的なレベルにそろえます。" },
fadeIn: { type: 'number', description: "フェードインで、単位は秒です。" },
fadeOut: { type: 'number', description: "フェードアウトで、単位は秒です。" },
}}
/>

```ts
const { jobId } = await client.audio.adjustVolume({ audioUrl: musicUrl, volume: 60, fadeOut: 3 })
```

### audio.combine(input)
オーディオのセグメントを順番につなげます（`POST /v1/combine-audio`）。[**オーディオを結合**（Combine Audio）](https://nodaro.ai/docs/nodes/audio/combine-audio)ノードと同じ処理です。

```ts
combine(input: { segments: Array<{ url: string; startTime?: number; endTime?: number }> }): Promise<{ jobId: string }>
```

<TypeTable
type={{
segments: { type: 'Array<{ url: string; startTime?: number; endTime?: number }>', required: true, description: "順番に並んだセグメントです。それぞれ、ファイルの一部だけを使うこともでき、その場合は秒単位の startTime から endTime までです。" },
}}
/>

```ts
const { jobId } = await client.audio.combine({
segments: [{ url: introUrl }, { url: episodeUrl, startTime: 4 }, { url: outroUrl }],
})
```

### audio.transcribe(input)
オーディオまたは動画ファイルの音声をテキストにします（`POST /v1/transcribe`）。[**文字起こし**（Transcribe）](https://nodaro.ai/docs/nodes/audio/transcribe)ノードと同じ処理です。

```ts
transcribe(input: {
audioUrl: string
provider?: "elevenlabs-stt" | "incredibly-fast-whisper" | "whisper"
language?: string
diarize?: boolean
tagAudioEvents?: boolean
wordTimestamps?: boolean
}): Promise<{ jobId: string }>
```

<TypeTable
type={{
audioUrl: { type: 'string', required: true, description: "オーディオまたは動画ファイルです。" },
provider: { type: '"elevenlabs-stt" | "incredibly-fast-whisper" | "whisper"', default: '"whisper"', description: "エンジンです。下の表を参照してください。" },
language: { type: 'string', description: "言語を強制します。省略すると言語が自動検出されます。" },
diarize: { type: 'boolean', description: "各単語を話しているのが誰かをラベルで示します。elevenlabs-stt のみです。" },
tagAudioEvents: { type: 'boolean', description: "笑い声や拍手などの音にタグを付けます。elevenlabs-stt のみです。" },
wordTimestamps: { type: 'boolean', description: "単語ごとのタイミングを求めます。" },
}}
/>

| `provider` | 単語のタイミング | 備考 |
| --- | --- | --- |
| [`elevenlabs-stt`](https://nodaro.ai/docs/models/audio/elevenlabs-stt) | 常に返す | `diarize` と `tagAudioEvents` に対応する唯一のエンジンです。 |
| [`incredibly-fast-whisper`](https://nodaro.ai/docs/models/audio/incredibly-fast-whisper) | `wordTimestamps: true` のときだけ | このフラグがない場合、ジョブは成功して課金されますが、フレーズ単位のセグメントのみで、`words` は空のリストになります。 |
| [`whisper`](https://nodaro.ai/docs/models/audio/whisper) | 返さない | フレーズ単位のセグメントのみです。`wordTimestamps: true` を指定すると、クレジットが使われる前に `400 validation_error` で拒否されます。 |

`provider` を省略すると `whisper` が使われるため、`provider` を指定しない `wordTimestamps: true` も、同じ 400 になります。単語のタイミングが必要な場合は、`elevenlabs-stt` を指定するか、`wordTimestamps: true` を付けた `incredibly-fast-whisper` を指定してください。キネティックスタイルの字幕には、これが必要です。

```ts
const { jobId } = await client.audio.transcribe({ audioUrl: talkUrl, provider: "elevenlabs-stt" })
```

完了したジョブの `output_data` は `TranscribeJobOutput` です。

| フィールド | 単位 | 内容 |
| --- | --- | --- |
| `text` | | 文字起こし全体を、1 つの文字列にしたものです。 |
| `language` | | 検出された、または指定した言語コードです。 |
| `words` | ミリ秒 | 単語ごとに 1 件で、`{ text, startMs, endMs, speaker? }` の形です。エンジンに単語のタイミングを求めなかった場合は空です。 |
| `json` | ミリ秒 | 正規化された文字起こしで、`{ version, language, words, segments? }` の形です。[`client.edit`](https://nodaro.ai/docs/developers/sdk/editing) が受け取る形式です。 |
| `segments` | **秒** | 元のフレーズの範囲で、古いエンジンのみが返します。`elevenlabs-stt` は返さないので、`words` を読み取ってください。 |

`words` は、[`client.media.addCaptions()`](https://nodaro.ai/docs/developers/sdk/media-and-uploads#mediaaddcaptionsinput) が受け取る `captions` と、まったく同じ形です。そのため、文字起こしを修正して、そのまま字幕として焼き込めます。

## Frequently asked questions

### Nodaro の SDK で、動画の声を変更するにはどうすればよいですか？

client.voices.change を、videoUrl と voiceId とともに呼び出します。サーバーが声を置き換え、新しい声を元の動画に載せ直します。ジョブをポーリングして output_data.videoUrl を取得してください。

### 録音の中の話者ごとに、異なるボイスを割り当てるにはどうすればよいですか？

client.voices.recast を、検出された話者の順に 1 件ずつ対応する orderedVoices とともに呼び出します。自分の声のままにする話者には null を使います。Nodaro Cloud で実行されます。

### SDK で、ボイスをクローンすることはまだできますか？

いいえ。ボイスクローンは 2026 年 9 月に廃止され、createClone は現在、410 voice_cloning_retired で失敗します。既存のクローンは引き続き使えます。代わりに client.voices.design で新しいボイスを作成してください。

### 単語のタイミングを返す文字起こしエンジンはどれですか？

elevenlabs-stt は常に単語のタイミングを返します。incredibly-fast-whisper は、wordTimestamps を true に設定した場合にだけ返し、whisper は決して返しません。
