# Choosing a model

> Pick the right image, video, audio or text model for each job in Nodaro, from low-cost draft models to premium final renders, with live credit prices.

Source: https://nodaro.ai/docs/guides/choosing-models

Nodaro runs more than a hundred image, video, audio and text models. The right one depends on what you make and how much you want to spend. This guide groups the models by job and by price, so you can choose without reading every model page. Start from the quick picks, draft on a low-cost model, and switch to a premium model for the final render. The tables on this page are read from Nodaro's model catalog, so their prices stay current.

## Quick picks by task

Find your goal in the left column, then click a model for its settings, prices and prompt tips.

- best for typography / logos / text-heavy: Nano Banana Pro, GPT Image 2. Nano Banana Pro for diagrams / complex text; GPT Image 2 for logos and short copy.
- cheapest realistic image: Z-Image, Qwen, Imagen 4 Fast. Z-Image is the cheapest. Qwen / Imagen4 Fast for slightly higher quality.
- highest fidelity image: Nano Banana Pro, Imagen 4 Ultra, Flux 2 Flex. Pick by family preference; all three are premium tiers.
- image edit / restyle: Flux Kontext Pro, Ideogram Remix, Seedream 5 Pro (I2I). Flux Kontext preserves identity; Ideogram Remix is character-aware; Seedream 5 Pro for instruction-based edits (5 Lite is the budget option).
- highest-resolution image: Topaz Image Upscale, Nano Banana Pro, GPT Image 2. Generate at the model's top tier, then Topaz upscale 4x (Topaz's only lever is the 1x/2x/4x factor).
- background removal / cutout: Recraft Remove BG. Cheap, no prompt needed.
- best cinematic video: VEO 3.1 Quality, Kling 3.0, Seedance 2. VEO 3.1 Quality for premium narrative; Kling 3.0 for music-synced motion; Seedance 2 for reference-driven consistency.
- cheap batch video clips: VEO 3.1 Fast, Wan 2.2 Turbo, Bytedance Lite I2V. VEO 3.1 Fast is the best price/quality balance with native audio.
- video with start + end frame: VEO 3.1 Quality, VEO 3.1 Fast, Kling 2.5 Turbo Pro, Hailuo 02 I2V Pro, Hailuo 02 Standard, Seedance 2. All listed support an end frame; VEO uses imageUrls[start, end].
- music / song generation: Suno V6, Suno V6 Wild, Suno V6 Mini, Suno v5.5. V6 is the default flagship; V6 Wild for bolder, less predictable results; V6 Mini when speed matters; v5.5 / v5 / v4 keep their own character. Same price.
- voice over / narration: ElevenLabs v3, ElevenLabs Turbo v2.5. v3 supports [audio tags] for emotion; Turbo is cheaper for plain narration.
- lip-sync a portrait to audio: Kling Avatar Pro, Kling Avatar Standard, InfiniTalk. Pro for best mouth shape; InfiniTalk for resolution control.
- transcription / captions: ElevenLabs STT, Incredibly Fast Whisper, Whisper. Captions need WORD timestamps: ElevenLabs STT (always) or Incredibly Fast Whisper. Plain Whisper returns phrase segments only.
- motion transfer (drive a subject by another video): Kling 2.6 Motion Transfer, Kling 3.0 Motion Transfer. Kling 2.6 base is cheap; Kling 3.0 is premium.

## Everyday, standard and premium

Models fall into three price tiers. The tiers are relative within one kind of media.

| Tier | Use it for |
| --- | --- |
| **Everyday** | Low cost and fast. The right default for drafts, iteration and most work. |
| **Standard** | A step up in quality for a moderate price. |
| **Premium** | The best quality at the highest price. Use it for hero shots and final renders. |

**Compare credits, not tier names, across kinds of media.** A premium image model still costs far less than a premium video model.

**Prices are for the default settings.** The credits in each table below are the price of the model's default settings. A higher resolution, a longer clip or native audio costs more. See [Credits](https://nodaro.ai/docs/concepts/credits) for how prices are charged.

## Image models

| Job | Model | Why |
| --- | --- | --- |
| Detailed images, text in the image, diagrams | [Nano Banana Pro](https://nodaro.ai/docs/models/image/nano-banana-pro) | The default of Generate Image. Best for text rendering, diagrams and complex compositions. |
| Fast iteration and high-volume work | [GPT Image 2.5 Flare](https://nodaro.ai/docs/models/image/gpt-image-2-5-flare) | Higher quality than GPT Image 2 at about half the time. Good for social posts, campaign variants and thumbnails. |
| The final, brand-sensitive render | [GPT Image 2.5 Sunburst](https://nodaro.ai/docs/models/image/gpt-image-2-5-sunburst) | Takes longer for tighter control and detail: packaging, diagrams, product retouching. |
| The cheapest drafts and storyboards | [Z-Image](https://nodaro.ai/docs/models/image/z-image), [Nano Banana 2 Lite](https://nodaro.ai/docs/models/image/nano-banana-2-lite) | Z-Image is the cheapest model in the catalog, with a limited set of aspect ratios. |
| Following detailed instructions | [Seedream 5 Pro](https://nodaro.ai/docs/models/image/seedream-5-pro) | The strongest instruction following and visual reasoning of the Seedream family. |
| Editing while keeping the subject | [Flux Kontext Pro](https://nodaro.ai/docs/models/image/flux-kontext-pro) | Preserves the subject's identity through edits and style changes. |
| Logos, short copy and stylized illustration | [GPT Image 2](https://nodaro.ai/docs/models/image/gpt-image-2), [Ideogram V3](https://nodaro.ai/docs/models/image/ideogram-v3) | Strong typography. |
| The highest resolution | [Topaz Image Upscale](https://nodaro.ai/docs/models/image/topaz-image-upscale) | Generate at the model's top resolution, then upscale the image 2x or 4x. |
| Background removal | [Recraft Remove BG](https://nodaro.ai/docs/models/image/recraft-remove-bg) | Low cost, and no prompt needed. |

If you are unsure between the two GPT Image 2.5 models, draft on Flare and finish on Sunburst. They cost the same.

| Model | Maker | Modes | Credits | Details |
| --- | --- | --- | --- | --- |
| [Z-Image](https://nodaro.ai/docs/models/image/z-image) | Tongyi-MAI | Text to image | 2 | Cheapest model in catalog. Fast, stylized output. Limited aspect ratios. |
| [Nano Banana 2 Lite](https://nodaro.ai/docs/models/image/nano-banana-2-lite) | Google | Text to image, Image to image | 10 | Lightweight Nano Banana 2 (Gemini 3.1 Flash-Lite) — fast, low-cost 1K generation and editing. |
| [GPT Image 2.5 Flare](https://nodaro.ai/docs/models/image/gpt-image-2-5-flare) | OpenAI | Text to image | from 15 | Fast everyday GPT Image 2.5 - higher quality than GPT Image 2 at about half the latency. The default of the pair: social and creator content, campaign variants, thumbnails, rapid iteration, high-volume work. |
| [GPT Image 2.5 Sunburst](https://nodaro.ai/docs/models/image/gpt-image-2-5-sunburst) | OpenAI | Text to image | from 15 | Precision GPT Image 2.5 - trades generation time for tighter control and detail fidelity. Pick it for brand-sensitive and production work: packaging, diagrams, ecommerce retouching, polished campaign creative. |
| [Nano Banana 2](https://nodaro.ai/docs/models/image/nano-banana-2) | Google | Text to image, Image to image | from 20 | Newer Nano Banana with native resolution control (1K/2K/4K) and Google Search context. |
| [Nano Banana Pro](https://nodaro.ai/docs/models/image/nano-banana-pro) | Google | Text to image, Image to image | from 45 | Top-tier Nano Banana — best for text rendering, diagrams, and complex compositions. |
| [Seedream 5 Pro](https://nodaro.ai/docs/models/image/seedream-5-pro) | Bytedance | Text to image | from 18 | Flagship Seedream 5 Pro — strongest instruction following and visual reasoning. Basic = 1K, high = 2K. |
| [Flux Kontext Pro](https://nodaro.ai/docs/models/image/flux-kontext-pro) | Black Forest Labs | Text to image, Image editing | 13 | Context-aware editing and style transfer. Strong at preserving subject identity through edits. |
| [Ideogram V3](https://nodaro.ai/docs/models/image/ideogram-v3) | Ideogram | Text to image | 18 | Strong typography and stylized illustration. Speed/quality tiered (TURBO/BALANCED/QUALITY). |
| [Topaz Image Upscale](https://nodaro.ai/docs/models/image/topaz-image-upscale) | Topaz | Image upscaling | from 25 | High-quality image upscale at 1x (enhance only), 2x or 4x. Best for production-ready output. |
| [Recraft Remove BG](https://nodaro.ai/docs/models/image/recraft-remove-bg) | Recraft | Background removal | 3 | Remove image background. Cheap utility. |

Every image model is on [Image models](https://nodaro.ai/docs/models/image).

## Video models

Before you choose, check three things on the model page. Does the model make **native audio**? Which **durations** does it offer? Does it accept an **end frame** as well as a start frame? Many video models make both text-to-video and image-to-video under one name. [Generate Video](https://nodaro.ai/docs/nodes/video/generate-video) chooses the mode from what you connect: with a start frame it animates the image, without one it works from the prompt.

| Job | Model | Why |
| --- | --- | --- |
| Most clips | [VEO 3.1 Fast](https://nodaro.ai/docs/models/video/veo-3-1-fast) | The best balance of price and quality, with native audio, in 4, 6 or 8 seconds. |
| Many clips at the lowest price | [VEO 3.1 Lite](https://nodaro.ai/docs/models/video/veo-3-1-lite), [Wan 2.2 Turbo](https://nodaro.ai/docs/models/video/wan-2-2-turbo), [Bytedance Lite I2V](https://nodaro.ai/docs/models/video/bytedance-lite-i2v) | Built for high volume, such as B-roll. |
| Premium, cinematic shots | [VEO 3.1 Quality](https://nodaro.ai/docs/models/video/veo-3-1-quality) | Premium narrative video with native audio and an optional end frame. |
| Motion synced to music | [Kling 3.0](https://nodaro.ai/docs/models/video/kling-3-0) | Premium motion, 3 to 15 seconds, with native audio. |
| Consistency from reference images | [Seedance 2](https://nodaro.ai/docs/models/video/seedance-2) | Reference-driven video with native audio. [Seedance 2 Mini](https://nodaro.ai/docs/models/video/seedance-2-mini) is the budget tier. |
| One long shot | [Seedance 2.5](https://nodaro.ai/docs/models/video/seedance-2-5), [Wan 3.0](https://nodaro.ai/docs/models/video/wan-3-0) | Up to 30 seconds in one shot. |
| A start frame and an end frame | [VEO 3.1 Quality](https://nodaro.ai/docs/models/video/veo-3-1-quality), [VEO 3.1 Fast](https://nodaro.ai/docs/models/video/veo-3-1-fast), [Kling 2.5 Turbo Pro](https://nodaro.ai/docs/models/video/kling-2-5-turbo-pro), [Hailuo 02 I2V Pro](https://nodaro.ai/docs/models/video/hailuo-02-i2v-pro), [Seedance 2](https://nodaro.ai/docs/models/video/seedance-2) | Each supports an end frame. |
| A talking portrait | [Kling Avatar Pro](https://nodaro.ai/docs/models/video/kling-avatar-pro), [Kling Avatar Standard](https://nodaro.ai/docs/models/video/kling-avatar-standard), [InfiniTalk](https://nodaro.ai/docs/models/video/infinitalk) | Kling Avatar Pro gives the best mouth shapes. InfiniTalk lets you choose the resolution. |
| Motion from another video | [Kling 2.6 Motion Transfer](https://nodaro.ai/docs/models/video/kling-2-6-motion-transfer), [Kling 3.0 Motion Transfer](https://nodaro.ai/docs/models/video/kling-3-0-motion-transfer) | Kling 2.6 is the low-cost option. Kling 3.0 is premium. |
| A sharper final video | [Topaz Video Upscale](https://nodaro.ai/docs/models/video/topaz-video-upscale) | High-quality upscale and enhancement. |

| Model | Maker | Modes | Credits | Details |
| --- | --- | --- | --- | --- |
| [VEO 3.1 Lite](https://nodaro.ai/docs/models/video/veo-3-1-lite) | Google | Image to video, Text to video | from 75 | VEO 3.1 Lite — most cost-effective VEO tier for high-volume generation. 4/6/8s with audio, supports first+last frame. |
| [VEO 3.1 Fast](https://nodaro.ai/docs/models/video/veo-3-1-fast) | Google | Image to video, Text to video | from 150 | VEO 3.1 Fast — cheaper VEO 3.1 tier, 4/6/8s with audio. Good balance for most uses. Flat per-generation pricing across durations. |
| [VEO 3.1 Quality](https://nodaro.ai/docs/models/video/veo-3-1-quality) | Google | Image to video, Text to video | from 930 | Google VEO 3.1 Quality — premium cinematic video. 4/6/8s clips, optional end frame, native audio. No reference-to-video mode (Fast/Lite only). Flat per-generation pricing across durations. |
| [Kling 3.0](https://nodaro.ai/docs/models/video/kling-3-0) | Kuaishou | Image to video, Text to video | from 270 | Premium Kling 3.0 — variable 3-15s duration, native audio, 720P/1080P. |
| [Seedance 2 Mini](https://nodaro.ai/docs/models/video/seedance-2-mini) | Bytedance | Image to video, Text to video | from 120 | Budget Seedance 2 tier — 480p/720p only, per-second pricing by resolution. |
| [Seedance 2](https://nodaro.ai/docs/models/video/seedance-2) | Bytedance | Image to video, Text to video | from 230 | Seedance 2 — premium tier with native audio. Per-second pricing by resolution. |
| [Seedance 2.5](https://nodaro.ai/docs/models/video/seedance-2-5) | Bytedance | Image to video, Text to video | from 340 | Seedance 2.5 — up to 30s in one shot, native audio, wide multimodal references. 480p/720p/1080p. |
| [Wan 2.2 Turbo](https://nodaro.ai/docs/models/video/wan-2-2-turbo) | Alibaba | Image to video, Text to video | from 100 | Cheap, fast Wan turbo — 5s. Serves both i2v and t2v under one id. |
| [Bytedance Lite I2V](https://nodaro.ai/docs/models/video/bytedance-lite-i2v) | Bytedance | Image to video, Text to video | 57 | Cheapest Bytedance video tier with end-frame support. |
| [Kling 2.5 Turbo Pro](https://nodaro.ai/docs/models/video/kling-2-5-turbo-pro) | Kuaishou | Image to video, Text to video | from 110 | Faster Kling — good quality at lower cost. Supports end frame. |

Every video model is on [Video models](https://nodaro.ai/docs/models/video). For how two of them, Seedance 2.5 and Seedance 2 Fast, follow camera directions, read [Camera Motion Lab](https://nodaro.ai/docs/research/camera-motion-lab).

## Audio, voice and music models

| Job | Model | Why |
| --- | --- | --- |
| Voice-over with emotion | [ElevenLabs v3](https://nodaro.ai/docs/models/audio/elevenlabs-v3) | Supports audio tags in brackets for emotion and pacing. |
| Plain narration at a lower price | [ElevenLabs Turbo v2.5](https://nodaro.ai/docs/models/audio/elevenlabs-turbo-v2-5) | Fast and cheaper. |
| Speech in many languages | [ElevenLabs Multilingual v2](https://nodaro.ai/docs/models/audio/elevenlabs-multilingual-v2) | Built for multiple languages. |
| A conversation between several voices | [ElevenLabs Dialogue v3](https://nodaro.ai/docs/models/audio/elevenlabs-dialogue-v3) | Voices each role of a script. |
| Sound effects | [ElevenLabs Sound Effects](https://nodaro.ai/docs/models/audio/elevenlabs-sound-effects) | Short effects from a text prompt, at a very low price. |
| Songs and music | [Suno V6](https://nodaro.ai/docs/models/audio/suno-v6) | The default. [Suno V6 Wild](https://nodaro.ai/docs/models/audio/suno-v6-wild) is bolder and less predictable, [Suno V6 Mini](https://nodaro.ai/docs/models/audio/suno-v6-mini) is faster. All Suno models cost the same. |
| Transcripts for captions | [ElevenLabs STT](https://nodaro.ai/docs/models/audio/elevenlabs-stt), [Incredibly Fast Whisper](https://nodaro.ai/docs/models/audio/incredibly-fast-whisper) | Captions need a time for every word, and these return it. [Whisper](https://nodaro.ai/docs/models/audio/whisper) returns phrase timing only. |

| Model | Maker | Modes | Credits | Details |
| --- | --- | --- | --- | --- |
| [ElevenLabs v3](https://nodaro.ai/docs/models/audio/elevenlabs-v3) | ElevenLabs | Text to speech | 30 | Latest ElevenLabs TTS — supports [audio tags] for emotion / pacing. Direct API. |
| [ElevenLabs Turbo v2.5](https://nodaro.ai/docs/models/audio/elevenlabs-turbo-v2-5) | ElevenLabs | Text to speech | 15 | Fast, cheap ElevenLabs TTS via the direct ElevenLabs API. Good for narration. |
| [ElevenLabs Multilingual v2](https://nodaro.ai/docs/models/audio/elevenlabs-multilingual-v2) | ElevenLabs | Text to speech | 30 | Multi-language ElevenLabs TTS via the direct ElevenLabs API. |
| [ElevenLabs Dialogue v3](https://nodaro.ai/docs/models/audio/elevenlabs-dialogue-v3) | ElevenLabs | Multi-speaker dialogue | 25 | Multi-speaker dialogue via the direct ElevenLabs API — give it a script, it voices each role (any voice: premade, library, or cloned). |
| [ElevenLabs Sound Effects](https://nodaro.ai/docs/models/audio/elevenlabs-sound-effects) | ElevenLabs | Sound effects | 3 | Generate short sound effects from a text prompt. |
| [ElevenLabs STT](https://nodaro.ai/docs/models/audio/elevenlabs-stt) | ElevenLabs | Speech to text | 22 | Speech-to-text with WORD-level timestamps (always on), speaker diarization and audio-event tags. The engine to use when the transcript feeds captions. |
| [Incredibly Fast Whisper](https://nodaro.ai/docs/models/audio/incredibly-fast-whisper) | OpenAI | Speech to text | 40 | Fast Whisper speech-to-text. Returns WORD-level timestamps when asked, so its transcript can feed captions. |
| [Whisper](https://nodaro.ai/docs/models/audio/whisper) | OpenAI | Speech to text | 40 | Whisper speech-to-text — PHRASE-level segments only, NO word timestamps. Fine for a transcript or a static subtitle; not for word-timed (kinetic) captions. |
| [Suno V6](https://nodaro.ai/docs/models/audio/suno-v6) | Suno | Music | 30 | Suno V6 — greater musical expression with more natural vocals and richer details. The flagship and the default. |
| [Suno V6 Wild](https://nodaro.ai/docs/models/audio/suno-v6-wild) | Suno | Music | 30 | Suno V6 Wild — pushes creative boundaries for bolder, more distinctive musical expression; more varied, less predictable results. |
| [Suno V6 Mini](https://nodaro.ai/docs/models/audio/suno-v6-mini) | Suno | Music | 30 | Suno V6 Mini — lightweight and fast, balancing quality and speed for effortless creation. |

Every audio model is on [Audio models](https://nodaro.ai/docs/models/audio).

## Text models

Nodes that write text, such as [Prompt](https://nodaro.ai/docs/nodes/automate/prompt), [Generate Script](https://nodaro.ai/docs/nodes/video/generate-script) and [QA Check](https://nodaro.ai/docs/nodes/automate/qa-check), offer language models in three price tiers.

| Tier | Use it for | Relative price |
| --- | --- | --- |
| **Economy** | High-volume, simple text: fast drafts, short prompts, checks. | Lowest |
| **Standard** | Most writing and structured output. The balance of quality and price. | Middle |
| **Premium** | The hardest reasoning: long scripts, detailed scene plans, careful rewrites. | Highest |

The models on offer differ by node. The node's settings panel shows each option and its credit price.

## Draft cheap, finish premium

### Draft on an everyday model

Build the workflow on a low-cost model at a low resolution or a short duration. Run it until the composition, the prompt and the pickers are right.

### Keep the model in one node

Drive similar shots from a [List](https://nodaro.ai/docs/nodes/automate/list) into one generation node instead of copying the node for every shot. The node runs once per list item, so moving from the draft model to the final one is a single change.

### Render the final version

Switch to a premium model or a higher resolution, and run again. Compare the result with the draft before you spend on the rest of the workflow.

Workflow: A List of three scene prompts drives one Generate Image node, so a draft run and a final run differ by one model change.

- List → Generate Image (prompt)

## Tips

- **Use the right model for text in images.** For a sign, a label or a poster with words, use Nano Banana Pro or GPT Image 2.
- **Check the video features first.** A model without native audio needs a separate voice or music track. A model without end-frame support cannot land on a planned last frame.
- **Save money on retries.** Drafts on an everyday model cost a fraction of a premium run, so do your trial and error there.

## Frequently asked questions

### Which image model should I start with?

Start with Nano Banana Pro, the default of Generate Image, for detailed images and readable text. Use GPT Image 2.5 Flare to iterate fast and GPT Image 2.5 Sunburst for the final, brand-sensitive render. Use Z-Image for the cheapest drafts.

### Which video model gives the best results?

VEO 3.1 Quality for premium, cinematic shots, Kling 3.0 for music-synced motion, and Seedance 2 for consistency driven by reference images. VEO 3.1 Fast is the best balance of price and quality for most clips, with native audio.

### Which model should I use to make captions?

Use ElevenLabs STT or Incredibly Fast Whisper, because captions need a time for every word. Plain Whisper returns only phrase-level timing, which is fine for a transcript but not for word-by-word captions.

### Why is a model's price shown as a starting price?

The number in the tables is the price of the model's default settings. A higher resolution, a longer clip or native audio costs more. Every model page lists all of its prices.

### Can I switch the model of many nodes at once?

Not in one step, because the model is a setting of each generation node. To keep it to one change, drive similar shots from a List into one generation node. The node runs once per item, all with the same model.
