Nodaro Docs
DocumentationNode ReferenceModelsAI Agents (MCP)DevelopersSelf-hostingResearch
Video

Generate Video

Make a video clip from a prompt, a start frame or references with Seedance, VEO, Kling, Wan and more, with sound, end frames and per-model credit rules.

The Generate Video node turns a prompt, an image or a set of references into a video clip. It is the main video node in Nodaro: the same node makes text-to-video, image-to-video, first-and-last-frame and reference-driven clips, and it chooses the mode from what you connect. You choose one of the video models, and the result can feed any other video node, for example Combine Videos.

When to use it

  • You need a short clip from a description: B-roll, an establishing shot, a product spin, a social post.
  • You have a still image and want it to move. Connect it to Start Frame, for example a result of Generate Image.
  • You want a clip that starts on one picture and ends on another. Connect both Start Frame and End Frame.
  • You want the same character, product or place in many clips. Connect asset nodes or reference images.
  • You want to change an existing clip by instruction. Seedance 2.5 and Gemini Omni can edit a connected video.

For a clip longer than one model run allows, use Generate Video Pro. To continue a clip you already have, use Extend Video.

Quick start

Add the node

Press Tab on the canvas, or open the node picker in the sidebar, and choose Video › Create › Generate Video.

Write the prompt

Type the prompt in the node, or connect a Text node to the Prompt input. Describe what moves, how it moves and what the camera does.

Connect a start frame, if you have one

Connect an image node to Start Frame to animate that image. Leave it empty to make the clip from the prompt alone.

Choose the model and the clip

Open the settings panel. Choose a model under Provider, then the Duration (seconds), Resolution and Aspect Ratio. The choices change with the model. The same controls also appear in the strip under the node when you point at it.

Run it

Click Run. The clip appears on the node, and every earlier result stays in its result strip.

promptstart frameassetslookTextYour promptGenerate ImageOpening frameCharacter AssetMayaCamera MotionSlow push-inGenerate VideoSeedance 2Combine Videos
A prompt, an opening frame, a character and a camera-motion picker feed Generate Video; the clip then goes to Combine Videos.

How the node chooses a mode

There is no mode switch. When the node runs, it looks at what is connected and picks the mode:

What is connectedWhat the node makes
A prompt onlyA text-to-video clip.
Start FrameAn image-to-video clip that opens on your image.
Start Frame and End FrameA clip that moves from the first image to the last one, on models that support an end frame.
Image Refs, Video Refs or Audio RefsA reference-driven clip. The model uses the references to shape the result, on models that accept them.
A clip in Video Refs, on Seedance 2.5 or Gemini OmniAn edit of that clip: always on Gemini Omni, and on Seedance 2.5 when the prompt asks for a change. See Edit a clip.

Every model in the list can be chosen in every mode. A few things work differently on some models:

  • Some models need a start image. Kling 2.1 Master, Kling 3 Omni, Hailuo 2.3, Hailuo 2.3 Pro, Bytedance Pro Fast, HappyHorse 1.1 Ref2V and Grok Imagine Video 1.5 have no text-to-video mode. Run stays disabled until an image is connected to Start Frame or End Frame, and the button's tooltip names the model. Reference images alone do not count as a start frame.
  • An end frame alone becomes the start image. If you connect End Frame without Start Frame, the node sends that image as the image to animate.
  • One entry, two versions. Grok Imagine, Wan 2.6, Wan 2.7 and HappyHorse 1.1 each appear once in the model list. The node uses the text-to-video or the image-to-video version of the model, depending on whether an image is connected.

Inputs

Generate Video has eleven inputs on its left edge. Each input accepts only the kinds of nodes listed, and the editor rejects a connection that does not fit. An input that the chosen model cannot use is shown disabled. Click an input to see what is connected to it, jump to a connected node, disconnect it, or add a new compatible node.

InputAcceptsWhat it does
PromptText nodes, such as Text, Prompt, Generate Script and Combine Text, and pickersThe description of the clip.
NegativeText nodesWhat the model should avoid. Some models receive it as a real negative prompt; for the others, Nodaro adds it to the prompt as an "Avoid:" line.
Start FrameImage nodes, such as Upload Image, Generate Image and Modify ImageThe first frame of the clip.
End FrameImage nodesThe last frame of the clip, on models that support one.
Image RefsImage nodes, several at onceReference images for models that accept them. The order matters: drag them in the settings panel to reorder.
Video RefsVideo nodes, several at onceReference clips for Seedance 2, Seedance 2.5, Hailuo 3, Wan 3.0 and Gemini Omni. On Seedance 2.5 and Gemini Omni, a clip here can also be edited.
AudioAudio nodesA soundtrack for the finished clip. On LTX 2.3 Pro without a start frame, the audio drives an audio-to-video clip instead.
Audio RefsAudio nodes, several at onceAudio that shapes the generation itself, on the Seedance 2 family, Hailuo 3 and Wan 3.0. With a voice line, Seedance 2 and Hailuo 3 sync the speaker's lips to it.
AssetsCharacter, Location, Object, Animal/Creature and Create Face nodesLocked identities. Their approved pictures and descriptions go with the prompt, and you can mention them with @ in the prompt.
LookCamera and look pickers, such as Camera Motion, Lens, Lighting, Framing, Mood, Style, Temporal and TransitionEach picker adds its wording to the prompt.
ElementsSubject pickers, such as Person, Pose, Animal, Vehicle, Styling and Held PropEach picker adds its wording to the prompt.

The output, Video, is the URL of the generated clip. It can feed any number of nodes at once.

Audio and Audio Refs do different jobs. Audio is laid over the finished clip as its soundtrack. Audio Refs goes into the model and changes what it generates.

Settings

These settings appear for most models. The settings panel shows only the controls the chosen model supports.

SettingWhat it does
ProviderThe video model. New nodes start on Seedance 2.0 Fast for 5 seconds. Point at a model in the list to see what it supports.
PromptThe text of the prompt, when nothing is connected to the Prompt input. Each model has its own limit, from 1,000 characters on Kling 2.6 to 30,000 on Seedance 2.5. The editor counts down to your model's limit, and text past the limit is cut off.
Negative PromptWhat to avoid.
Duration (seconds)The length of the clip. The choices depend on the model. On the Seedance 2 family, Auto lets the model choose the length.
ResolutionThe output resolution, on models that offer a choice.
Aspect RatioThe shape of the clip. Adaptive or Auto (from image) matches the connected image.
Generate AudioWhether the model makes a soundtrack, on models where sound can be switched.
Frame fitHow a connected start or end frame is reshaped before the model sees it. See Start and end frames.
Send frames asWhether a connected frame is sent as a real frame or as a reference image. See Start and end frames.
Loop trimCuts the finished clip at its best loop point. See Make a clean loop.
Motion hint (injected into prompt)Adds a motion level to the prompt: Subtle, Moderate or Dynamic.
Save clip to "…" as reference videoAppears when a Character is connected. Saves the finished clip to that character's reference videos, under the Variant label you type, so you can reuse it as a video reference.
Pre & post textText that is always added before and after the prompt. It is hidden from people who use your workflow as an app. See Prompt pre and post text.
The Generate Video settings panel with Seedance 2.0 selected, its mode note, the prompt and the negative prompt.The Generate Video settings panel with Seedance 2.0 selected, its mode note, the prompt and the negative prompt.

Settings for specific models

ModelExtra settings
VEO 3.1 Quality, Fast and LiteGeneration Mode: Frame-to-Frame (the default) or Reference Mode, which uses 1 to 3 reference images instead of frames. Seed (optional), from 10000 to 99999. Generate Audio, on by default. Auto-translate prompt to English, on by default; turn it off to keep the prompt word for word.
Seedance 2, Seedance 2 Fast, Seedance 2 Mini and Seedance 2.5Generate Audio (default on), Enable Web Search and NSFW Content Filter. A note at the top of the panel shows the mode the connected inputs produce.
Hailuo 3 (minimax-h3)Resolution: 2K (the default) or 768P, which costs less per second. Audio is always generated.
Wan 3.0 and Wan 3.0 PrimeGenerate Audio (default on) for ambient sound. New settings start at 5 seconds, 720p and Adaptive.
Kling 2.6Enable Sound. A clip with sound costs more.
Kling 2.5 Turbo Pro and Kling 2.1 MasterCFG Scale, from 0 (creative) to 1 (strict prompt adherence).
Kling 3 OmniQuality: Standard (720p) or Pro (1080p), and Generate Audio.
Grok ImagineResolution (480p or 720p) and Mode: Normal, Fun or Spicy.
Bytedance Lite and Bytedance ProCamera Fixed and Seed (-1 for random).
Hailuo 2.3 and Hailuo 2.3 ProResolution: 768P or 1080P. At 1080P the clip can be 6 seconds only.
Gemini Omni and Gemini Omni FlashSource clip trim (seconds, ≤10s window) when a video is connected to Video Refs.

Kling 3.0 settings

When you choose Kling 3.0, the settings panel changes to three tabs: Scene, Shots and Elements.

  • Scene holds the model, the Mode (Pro (1080p), Standard (720p) or 4K (Ultra HD)), the Aspect Ratio, Sound Effects for lip-synced dialogue and sound, and the duration from 3 to 15 seconds.
  • Shots has Multi-Shot Mode, which splits the clip into 2 to 6 shots, each with its own prompt and length, up to 15 seconds in total. Sound is required in multi-shot mode, and an end frame is not supported.
  • Elements lets you define named characters or objects with a description and 2 to 4 reference images, then mention them in the prompt as @name.

Which model to choose

Generate Video can run every model below. Click a model for its settings, credit prices and prompt tips.

ModelMakerModesCreditsDetails
Hailuo 02 I2V ProMiniMaxImage to video, Text to video143Hailuo 02 Pro — strong photoreal motion, fixed 5-second clips. Supports end frame.
minimax-h3MiniMaxImage to video, Text to videofrom 230MiniMax Hailuo 3 — premium multimodal tier: first/last frame + image/video/audio references, native audio, 2K (default) or 768P output, 4-15s per-second pricing.
Hailuo 2.3 ProMiniMaxImage to videofrom 130Hailuo 2.3 Pro — newer Hailuo with 768P / 1080P resolutions.
Hailuo 2.3 StandardMiniMaxImage to videofrom 75Cheaper Hailuo 2.3 tier — good baseline quality.
Hailuo 02 StandardMiniMaxImage to video, Text to videofrom 75Hailuo 02 Standard — economical option with end-frame support.
VEO 3.1 QualityGoogleImage to video, Text to videofrom 930Google VEO 3.1 Quality — premium cinematic video. 4/6/8s clips, optional end frame, native audio. No reference-to-video mode (Fast/Lite only). Flat per-generation pricing across durations.
VEO 3.1 FastGoogleImage to video, Text to videofrom 150VEO 3.1 Fast — cheaper VEO 3.1 tier, 4/6/8s with audio. Good balance for most uses. Flat per-generation pricing across durations.
VEO 3.1 LiteGoogleImage to video, Text to videofrom 75VEO 3.1 Lite — most cost-effective VEO tier for high-volume generation. 4/6/8s with audio, supports first+last frame.
Gemini OmniGoogleImage to video, Text to videofrom 230Google multimodal video with native audio; text/image-to-video + video-edit.
Gemini Omni FlashGoogleImage to video, Text to videofrom 160Google Gemini Omni Flash — faster/cheaper Omni tier: multimodal video with native audio, text/image-to-video + video-edit.
Kling 2.6KuaishouImage to video, Text to videofrom 138Kling 2.6 I2V — strong motion realism. 5s/10s, optional native audio.
Kling 2.5 Turbo ProKuaishouImage to video, Text to videofrom 110Faster Kling — good quality at lower cost. Supports end frame.
Kling 3.0KuaishouImage to video, Text to videofrom 270Premium Kling 3.0 — variable 3-15s duration, native audio, 720P/1080P.
Kling 2.1 MasterKuaishouImage to videofrom 400Master tier I2V — strong cinematic quality.
Kling 3 OmniKuaishouImage to videofrom 250Kling 3 Omni — 3-15s, 720p/1080p, end frame + reference images, native audio.
Grok Imagine (I2V)xAIImage to videofrom 50Grok image-to-video — stylized motion. Up to 15s.
Grok Imagine Video 1.5xAIImage to videofrom 295Grok Imagine 1.5 image-to-video — 1–15s, 480p/720p, per-second pricing. Requires an input image.
Seedance 2BytedanceImage to video, Text to videofrom 230Seedance 2 — premium tier with native audio. Per-second pricing by resolution.
Seedance 2 FastBytedanceImage to video, Text to videofrom 180Cheaper / quicker Seedance 2 tier.
Seedance 2 MiniBytedanceImage to video, Text to videofrom 120Budget Seedance 2 tier — 480p/720p only, per-second pricing by resolution.
Seedance 2.5BytedanceImage to video, Text to videofrom 340Seedance 2.5 — up to 30s in one shot, native audio, wide multimodal references. 480p/720p/1080p.
Bytedance Lite I2VBytedanceImage to video, Text to video57Cheapest Bytedance video tier with end-frame support.
Bytedance Pro I2VBytedanceImage to video, Text to video175Pro Bytedance video tier — better quality.
Bytedance Pro Fast I2VBytedanceImage to video90Faster Bytedance Pro variant.
Wan 3.0AlibabaImage to video, Text to videofrom 160Wan 3.0 — multimodal: first/last frame or image/video/audio references, native audio, 2-30s at 480p/720p/1080p.
Wan 3.0 PrimeAlibabaImage to video, Text to videofrom 250Wan 3.0 Prime — Alibaba's high-speed Wan 3.0 tier: same multimodal surface and 2-30s range, faster turnaround at a higher per-second rate.
Wan 2.6 I2VAlibabaImage to videofrom 175Wan 2.6 image-to-video — 5/10/15s at 720p/1080p.
Wan 2.2 TurboAlibabaImage to video, Text to videofrom 100Cheap, fast Wan turbo — 5s. Serves both i2v and t2v under one id.
Wan 2.6AlibabaVideo to video, Text to videofrom 175Wan 2.6 — text-to-video and video-to-video under a single id.
Wan 2.7 I2VAlibabaImage to video188Wan 2.7 image-to-video — 2–15s at 720p/1080p, supports start+end frame.
Wan 2.7 T2VAlibabaText to video188Wan 2.7 text-to-video — 2–15s at 720p/1080p.
LTX 2.3 ProLightricksImage to video, Text to videofrom 40Lightricks LTX 2.3 Pro — text/image/audio→video up to 4K, 6/8/10s, end-frame interpolation.
LTX 2.3 FastLightricksImage to video, Text to videofrom 180Lightricks LTX 2.3 Fast — text/image→video up to 20s at 1080p (6/8/10s at 2K and 4K). No audio input, no extend.
HappyHorse 1.1HappyHorseText to video282HappyHorse 1.1 text-to-video — 3–15s at 720p/1080p, 9 aspect ratios incl. 21:9/9:21, per-second pricing.
HappyHorse 1.1 I2VHappyHorseImage to video282HappyHorse 1.1 image-to-video — 3–15s at 720p/1080p, aspect ratio inferred from input image, per-second pricing.
HappyHorse 1.1 Ref2VHappyHorseImage to video282HappyHorse 1.1 reference-to-video — 1–9 reference images, 3–15s at 720p/1080p, per-second pricing.
Runway Gen-3RunwayImage to video, Text to video30Runway Gen-3. 5/10s at 720p/1080p.
  • Seedance 2.0 Fast — the default. Low cost, sound on, references supported. Use it to iterate.
  • VEO 3.1 Fast — a good balance of price and quality, with native audio and the same price for 4, 6 or 8 seconds.
  • VEO 3.1 Quality and Kling 3.0 — premium cinematic shots. Kling 3.0 adds multi-shot clips and lengths from 3 to 15 seconds.
  • Seedance 2 and Seedance 2.5 — the widest reference support: images, videos and audio together. Seedance 2.5 makes up to 30 seconds in one run and can edit a clip.
  • Wan 3.0 — 2 to 30 seconds at any whole second, with references and ambient audio.
  • Hailuo 3 (minimax-h3) — premium multimodal clips at 2K, with references and audio always on.
  • VEO 3.1 Lite, Wan 2.2 Turbo and Bytedance Lite — the cheapest clips, for batches and drafts.

For a comparison of every model, read Choosing a model.

One start frame, three models

The three clips below start from the same image, which Generate Image made, and use the same prompt with sound on. Only the model changes.

Prompt
Slow dolly in toward the boat as it drifts, mist rolling across the water, gulls in the distance, gentle lapping waves
VEO 3.1 Fast, 720p, 6 seconds
Seedance 2.5, 480p, 8 seconds
Gemini Omni Flash, 720p, 6 seconds
  • VEO 3.1 Fast kept the golden light of the start frame, pushed in slowly and added the gulls from the prompt.
  • Seedance 2.5 stayed closest to the start frame, with the gentlest motion. At 480p it is also the softest picture, which is fine for a draft.
  • Gemini Omni Flash left the start frame after about a second and a half: the light turned cooler and the mountains changed.

When a clip has to continue its start frame exactly, watch the first seconds of each take before you build on it, and try a second model when the first one drifts.

Lengths and resolutions by model

The length and the resolution you can choose depend on the model. The model pages list every value, and the aspect ratios too.

ModelLengthsResolutionsEnd frameReferences
Seedance 24–15 s, or Auto480p, 720p, 1080p, 4KYes9 images, 3 videos, 3 audio
Seedance 2 Fast4–15 s, or Auto480p, 720pYes9 images, 3 videos, 3 audio
Seedance 2 Mini4–15 s, or Auto480p, 720pYes9 images, 3 videos, 3 audio
Seedance 2.54–30 s, or Auto480p, 720p, 1080pYes30 images, 10 videos, 10 audio
Hailuo 3 (minimax-h3)4–15 s2K, 768PYes9 images, 3 videos, 3 audio
VEO 3.1 Quality4, 6, 8 s720p, 1080p, 4KYesNone
VEO 3.1 Fast4, 6, 8 s720p, 1080p, 4KYesUp to 3 images
VEO 3.1 Lite4, 6, 8 s720p, 1080p, 4KYesUp to 3 images
Gemini Omni4, 6, 8, 10 s720p, 1080p, 4KNo7 inputs; a video counts as 2
Gemini Omni Flash4, 6, 8, 10 s720p, 1080p, 4KNo7 inputs; a video counts as 2
Kling 3.03–15 s720p, 1080p, 4KYesElements
Kling 3 Omni3–15 s720p, 1080pYesUp to 7 images
Kling 2.65, 10 s—NoNone
Kling 2.5 Turbo Pro5, 10 s—YesNone
Kling 2.1 Master5, 10 s—NoNone
Wan 3.0 and Wan 3.0 Prime2–30 s480p, 720p, 1080pYes10 images, 5 videos, 5 audio
Wan 2.72–15 s720p, 1080pYesNone
Wan 2.65, 10, 15 s720p, 1080pNoNone
Wan 2.2 Turbo5 s480p, 720pNoNone
HappyHorse 1.13–15 s720p, 1080pNoNone
HappyHorse 1.1 Ref2V3–15 s720p, 1080pNo1 to 9 images
Hailuo 2.3 Pro and Hailuo 2.3 Standard6, 10 s768P, 1080PNoNone
Hailuo 02 Standard6, 10 s512P, 768PYesNone
Hailuo 02 I2V Pro5 s—YesNone
Bytedance Lite I2V5, 10 s480p, 720p, 1080pYesNone
Bytedance Pro I2V5, 10 s480p, 720p, 1080pNoNone
Bytedance Pro Fast I2V5, 10 s720p, 1080pNoNone
Grok Imagine (I2V)6, 10 s480p, 720pNoReference images
Grok Imagine Video 1.51–15 s480p, 720pNoNone
LTX 2.3 Pro6, 8, 10 s1080p, 2K, 4KYesNone
LTX 2.3 Fast6–20 s at 1080p; 6, 8, 10 s at 2K and 4K1080p, 2K, 4KYesNone
Runway Gen-35, 10 s720p, 1080pNoNone

A dash means the model offers no choice of resolution.

Aspect ratios. VEO 3.1, Gemini Omni and LTX 2.3 make 16:9 or 9:16 clips. The Seedance 2 family adds 1:1, 4:3, 3:4, 21:9 and Adaptive, which is its default and matches the connected image. Hailuo 3 and Wan 3.0 also offer Adaptive as their default. A pure text-to-video clip on Hailuo 3 or Wan 3.0 renders 16:9 unless you choose a ratio. HappyHorse 1.1 offers nine ratios for text-to-video and Ref2V clips, including 4:5, 5:4, 21:9 and 9:21; for image-to-video it follows the input image.

4K on VEO 3.1. VEO makes the clip at 1080p, then upscales it to 4K in the same run, billed at the 4K price. 4K on Gemini Omni is not available on the free tier.

Start and end frames

Connect an image to Start Frame to open the clip on it. Connect a second image to End Frame to end the clip on it. These models accept an end frame:

  • VEO 3.1 Quality, Fast and Lite.
  • The Seedance 2 family and Seedance 2.5.
  • Hailuo 3, Hailuo 02 Standard and Hailuo 02 I2V Pro.
  • Kling 3.0, Kling 3 Omni and Kling 2.5 Turbo Pro.
  • Wan 2.7, Wan 3.0 and Wan 3.0 Prime.
  • Bytedance Lite.
  • LTX 2.3 Pro and LTX 2.3 Fast.

Other models ignore the end frame.

Frame fit

A frame that is not the exact size the model renders gets reshaped by the model, and some models do it one frame into the clip. The first frame is your image, and the rest of the clip is slightly taller or wider. Frame fit fixes the size before the frame is sent:

ValueWhat it does
Match output size (default)Resizes the frame to the exact pixel size the model renders for your resolution and aspect ratio.
Match aspect ratioCorrects only the aspect ratio, with the smallest possible change.
Keep originalSends the frame untouched.

When Nodaro knows the size for your model, resolution and ratio, the panel names it under the setting, for example 720×1280. When your image is more than 5% away from the target shape, for example a square photo in a 9:16 clip, the image is first cropped at the center to the target ratio, then resized. The subject is never squashed. For a combination Nodaro has not measured, the frame is left as it is.

Send frames as

ValueWhat it does
Auto (default)Picks the best way for the model. The option shows what it resolves to: Auto (as frame) or Auto (as reference image).
FrameAlways sends a real start frame.
Reference imageAlways sends the image as a reference, and the prompt names it as the opening frame.

The Seedance 2.0 family (Fast, standard and Mini) zooms a real start frame by about 2% and loses brightness within six frames. Sent as a reference image, the same models keep the look from the first frame. Every other model Nodaro measured reproduces the opening frame more faithfully as a real frame, so Auto keeps them in frame mode. On a model that takes no reference images, this setting is hidden.

Two limits apply. First, a frame never pushes out one of your own reference images: if the frames would exceed the model's image limit, the frame stays a frame. Second, when your request already has reference images, several models cannot keep a real start frame. Seedance and Wan move the frame into the reference list, Hailuo 3 switches to its reference mode, and VEO 3.1 drops the end frame.

References: images, videos and audio

References tell a model what to include without pinning an exact frame. Connect them to Image Refs, Video Refs and Audio Refs. The limits per model are in the table above. Models that do not support references ignore them.

  • Seedance 2 family and Hailuo 3. Frames and references can be connected together. With no reference, the frames are exact first and last frames. With any reference connected, the frames join the reference set and the prompt names them as the opening or closing frame. The panel shows the resolved mode and warns you when an input will be dropped.
  • Seedance 2.5. Up to 30 images, 10 videos and 10 audio clips. Each reference video must be 2 to 30 seconds long, and all of them together 30 seconds at most. Reference audio can be up to 30 seconds per clip.
  • Seedance 2 Fast. Each reference audio clip must be 15.2 seconds or shorter.
  • Hailuo 3. Each reference video must be 2 to 15 seconds, and 15 seconds in total. Reference audio is up to 15 seconds per clip, and it must come with an image or video reference; audio alone is not accepted.
  • Wan 3.0. Up to 10 images, 5 videos and 5 audio clips. Each video and audio clip is 1 to 15 seconds, with 15 seconds in total per kind. With a reference video, the reference and the output together must stay at 30 seconds or less. A start or end frame joins the reference list, after your own images.
  • VEO 3.1 Fast and Lite. Connecting reference images switches VEO to reference mode, with or without a start frame. Reference mode always makes an 8-second clip; VEO's price is the same for 4, 6 and 8 seconds, so the cost does not change. VEO takes at most three images: a start frame takes the first place, and references fill the rest. A connected end frame is dropped in this mode. VEO 3.1 Quality has no reference mode, so its reference inputs are disabled.
  • Gemini Omni and Gemini Omni Flash. Up to seven inputs: the start frame and each reference image count as one, a video counts as two. With a start frame, references act as identity references, not frames. Inputs over the limit are dropped, and the node says how many.
  • HappyHorse 1.1 Ref2V takes 1 to 9 reference images. Kling 3 Omni takes up to 7.

Most of these length limits are checked before the run starts, so a clip that breaks them costs no credits. Wan 3.0 checks its reference video limits during the run instead, and refunds the reserved credits when a clip is too long. Trim long clips first with Trim Video.

Point the prompt at a reference

On reference-capable models, you can tie a phrase in the prompt to one reference, so that it drives the result. Type a token where the phrase belongs. The @ autocomplete in the prompt editor offers them.

TokenBecomesUse
{image:1:person}the person from @image_1Tie a subject to reference image 1
{image:2:jacket}the jacket from @image_2Tie an item to reference image 2
{image:1}the subject in @image_1Tie without a noun
{video:1:clip}the clip from @video_1Tie to reference video 1
{audio:1:voice}the voice from @audio_1Tie to reference audio 1

For example, with two reference images connected, the prompt circle {image:1:person} wearing {image:2:jacket} for a 360 spin becomes "circle the person from @image_1 wearing the jacket from @image_2 for a 360 spin".

  • Numbering. Image numbers count Image Refs first, in order, then the assets on the Assets input. With one reference image and a connected Object, the object is {image:2}. Videos and audio have their own numbers.
  • Frames are numbered for you. Start and end frames come after your images, and the prompt names them as the opening or closing frame. You never write a token for a frame.
  • A number out of range becomes plain text. {image:5:ghost} on a node with two references becomes "ghost".
  • Models without references turn every token into its plain label, so the prompt still reads naturally.

Sound

Many models make an audio track together with the picture. The way you control it depends on the model:

  • VEO 3.1 makes sound from the prompt, on by default. Turn off Generate Audio for a silent clip.
  • The Seedance 2 family and Wan 3.0 make sound by default. Turn it off with Generate Audio (default on).
  • Hailuo 3 always makes sound. With reference audio connected, it syncs the lips to it.
  • Gemini Omni makes a soundtrack on every run and takes its direction from the prompt, for example "Sound design: gentle breeze, distant bird chirps". It cannot be switched off, and it does not speak dialogue.
  • Kling 2.6 makes sound only when Enable Sound is on. Kling 3.0 has Sound Effects. Kling speaks scripted dialogue when you quote the line in the prompt, for example [Anna: warm calm voice]: "good morning". Kling 2.6 speaks English and Chinese; other languages are translated to English by the model.
  • Kling 3 Omni includes audio in its price, with a Generate Audio switch.

To use your own soundtrack, connect it to the Audio input. To make the model react to a voice or music, connect it to Audio Refs instead.

Edit a clip

Seedance 2.5 and Gemini Omni can change a clip you connect to Video Refs.

With Seedance 2.5, the model decides from your prompt whether the run is a normal generation that uses the clip as a reference, or an edit of the clip, for example "remove the sign" or "make it black and white". An edit keeps the source clip's aspect ratio and length. If the model reads your prompt as an edit and your settings ask for another ratio or length, Nodaro submits the run once more with the values the edit needs, and you get the edited clip instead of an error. To ask for an edit from the start, set Aspect Ratio to Adaptive and Duration to Auto. The source clip must be 4 to 30 seconds long. The factory preset Edit Video sets all of this for you.

With Gemini Omni and Gemini Omni Flash, a connected clip switches the node to video edit. The clip is trimmed to a window of at most 10 seconds, which you set under Source clip trim. The duration follows the clip, and a video edit has a flat price per run.

Make a clean loop

Loop trim finds the frame where the clip loops back most cleanly and cuts the clip there. Turn it on in the settings panel:

  • Frames to test sets how many frames Nodaro compares, from 4 to 64. The default is 16.
  • Quality: Precise cuts at the exact frame with a slight quality loss. Lossless cuts on a keyframe and keeps the file byte for byte.

For a perfect loop, connect the same image to Start Frame and End Frame. Without an end frame, Nodaro still picks the best loop point it finds, but the loop may show a jump. Loop trim also removes the fade that VEO 3.1 adds at the end of a loop.

Loop trim adds a small charge: one credit per 5 seconds of clip, plus one credit per 24 frames tested, each rounded up. For example, an 8-second clip with 16 frames tested adds 3 credits. If the trim fails, you keep the untrimmed clip and the loop trim credits are refunded.

Keep a character consistent

  • Connect assets. Connect a Character Asset, Location Asset, Object/Props Asset or Animal/Creature Asset node to Assets. Nodaro sends their approved pictures and descriptions with the prompt.
  • Mention them. Write @maya in the prompt to place the character exactly where the sentence needs it. Give a mention a role, such as @maya:1:clothes, to say what to take from the reference. See Reference roles.
  • Check what is sent. The Injected references list in the settings panel shows every reference the model will receive. Drag to reorder, or click × to remove one.
  • Use a board. A reference board or a consistency grid from Generate Image keeps a face stable across clips. The Scene Recipes presets are built to read these boards. See Reference boards.

Read Consistent characters for the full method.

Presets

Open Presets on the node for ready-made setups. Each preset chooses a fitting model, aspect ratio and duration, fills the prompt with {placeholder || default} slots, and adds a tuned negative prompt against warping, flicker and morphing. You can run a preset unedited: every slot falls back to its default.

FolderExamples
Camera MovesSlow Push-In, Dolly Out, 360° Orbit, Crane Up, Tracking Follow, Whip Pan, Dolly Zoom (Vertigo)
Shot Types & AnglesEstablishing Wide, Close-Up, Macro, Low-Angle Hero, Overhead Top-Down, FPV Drone
Cinematic & SpecialtyHandheld Doc, Slow Motion, Timelapse, Hyperlapse, Bullet Time, Rack Focus
Social & ReelsVertical Hero, Talking Head, Product Reveal, Trend Quick-Cut, POV Walk, all at 9:16
Product & AdsProduct Hero, 360 Spin, Liquid Splash, Unboxing, Lifestyle Ad
Motion Graphics & LogoLogo Sting, Title Reveal, Particle Background, Loop Background
B-Roll & NatureClouds Timelapse, Water Slow-Mo, Forest Drift, Aerial Landscape, Ocean Loop
Animation & StyleAnime Motion, 3D Cartoon, Claymation, Living Watercolor
Looping & BackgroundsSubtle Motion, Living Wallpaper
Viral & EffectsFrozen in Ice, Superhero Transformation, Elevator Doors Reveal, POV Skydive, Underwater POV. These work best with an input image.
Scene RecipesViral Meteor Scene, Cartoon Short (Opening, Chase, Resolution), Two-Character Dialogue, Disaster Reveal, Chase Scene. Seedance 2 scenes with native audio and lip-synced quoted lines. Connect your reference boards or cast grids as reference images.
Video EditingEdit Video: edits a clip in Video Refs by instruction on Seedance 2.5. Type the change you want as the prompt.

To chain several Scene Recipes into one short, join the clips with the Seamless Join (One-Shot) preset of Combine Videos. Read Presets to save and share your own.

Credits

The cost badge on the node shows the price before you run. The price depends on the model, the duration and the resolution. On some models it also depends on sound and references.

  • Flat price per clip. VEO 3.1 costs the same for 4, 6 or 8 seconds. VEO 3.1 Lite costs 75 credits at 720p and 90 at 1080p. VEO 3.1 Fast costs 150 at 720p and 170 at 1080p. VEO 3.1 Quality costs 1,000 for a 1080p clip.
  • Price per duration. Gemini Omni and Gemini Omni Flash have a price per length and resolution. Gemini Omni Flash costs 160 credits for 4 seconds and 320 for 10 seconds at 720p or 1080p. A video edit has one flat price per run: 420 credits on Gemini Omni Flash and 600 on Gemini Omni, at 720p or 1080p.
  • Price per second. The Seedance 2 family, Seedance 2.5, Hailuo 3, Wan 3.0, Grok Imagine Video 1.5 and HappyHorse 1.1 are priced by the second. For example, Wan 3.0 costs 200 credits for 5 seconds at 720p and 640 for 8 seconds at 1080p.
  • References make Seedance cheaper. On the Seedance 2 family and Seedance 2.5, connecting any reference switches to a lower price per second. Seedance 2 at 1080p costs 2,040 credits for 8 seconds, or 1,240 with a reference.
  • Sound can cost more. Kling 3.0 costs 680 credits for 10 seconds without sound and 1,000 with sound. On Kling 3.0, sound is on unless you turn it off.
  • Reference videos count their own length. On the Seedance 2 family and Hailuo 3, the price covers the seconds of each reference video plus the seconds of output. Hailuo 3 at 2K with 8 seconds of output and a 5-second reference video costs 1,187 credits. On Wan 3.0, a reference video adds nothing.
  • Extra images on Hailuo 3. Each input image beyond the first five adds a small charge. Six seconds at 2K with eight images costs 630 credits instead of 550.
  • Auto length is settled after the run. With Duration set to Auto, Nodaro reserves the price of the model's longest clip: 30 seconds on Seedance 2.5 and 15 seconds on the other Seedance 2 models. When the clip is ready, you pay for the length you got and the rest is refunded. Your balance must cover the reservation for the run to start.
  • Edits are settled after the run too. A Seedance 2.5 run with a reference video is reserved for the longer of your duration and the clip, then settled to the length delivered.

Every model's page lists its exact prices. See Credits for how reservations and refunds work.

When a setting is not supported

A saved workflow can carry a resolution, aspect ratio or duration that the current model does not offer, for example after you switch models. Nodaro does not fail the run. It corrects the value to one the model accepts, and you pay for the corrected value, because that is also what the model makes.

  • The nearest value wins. A 4K request on a model that stops at 1080p renders 1080p. A portrait 9:21 becomes 9:16, never landscape.
  • An empty resolution is sent at the resolution it is priced at. The value used is recorded on the run.
  • Adaptive and Auto stay as they are. They tell the model to match the input.
  • Duration stays as you set it, except on LTX 2.3, which moves to its nearest length. A 7-second request on LTX 2.3 renders and bills 6 seconds.
  • Hailuo 3 and Wan 3.0 render a fixed default for an unsupported resolution: 2K on Hailuo 3 and 720p on Wan 3.0. The price follows what they render.

When a model blocks the prompt

Some models check a finished clip against their content policies, for example for resemblance to protected film and TV content, and reject it. When that happens, Nodaro rewrites the prompt once. The rewrite keeps the same subjects, camera language and mood, and softens what triggered the rejection. Nodaro then runs the clip again at no extra credit cost. If the second run is also rejected, the run fails with the model's own reason.

Older Text to Video and Image to Video nodes

Generate Video replaces the older Text to Video and Image to Video nodes. A saved workflow that contains them opens with Generate Video nodes in their place. The connections move to the matching inputs, and the models, settings and prices stay the same. You do not need to change anything.

Tips

  • Put the important words first. Text past your model's prompt limit is cut off, and Kling 2.6 allows only 1,000 characters.
  • Draft cheap, finish sharp. Iterate on Seedance 2.0 Fast or VEO 3.1 Lite at a low resolution, then switch the model or the resolution for the final clip.
  • Describe motion, not the picture. With a start frame, the image already shows the scene. Spend the prompt on what moves and on the camera.
  • Put the look in pickers. Pickers such as Camera Motion and Lighting add tested wording, and you can change them without rewriting the prompt.
  • Drive many shots from one node. Wire a List of shot prompts into one Generate Video node instead of copying the node, so the model is set once for every shot.
  • Connect only what the model uses. The settings panel hides controls the model does not support, but values you set for another model stay in the node.

Troubleshooting

Run is disabled and the tooltip names the model. The model has no text-to-video mode. Connect an image to Start Frame, or choose a model that makes clips from text.

The clip jumps slightly after the first frame. The model reshaped your start frame. Set Frame fit back to Match output size.

On Seedance 2.0, the first frames look zoomed or darker. Set Send frames as to Reference image, or leave it on Auto, which does this on the Seedance 2.0 family.

The node says an input will be dropped. You connected more references than the model accepts, or you combined inputs the model cannot use together, such as a VEO end frame with reference images. Remove inputs, or choose a model with a higher limit.

A reference clip is rejected before the run. A reference video or audio clip is outside the model's length limits. Trim it with Trim Video or Trim Audio, then run again. No credits are spent.

A Wan 3.0 run fails after it starts. Wan 3.0 checks its reference video limits while the run is in progress. The reserved credits are refunded. Shorten the reference videos and run again.

The Seedance 2.5 result has a different length or shape than you set. The model treated your prompt as an edit of the reference video, which keeps the clip's own ratio and length. See Edit a clip.

4K is rejected on Gemini Omni. 4K is not available on the free tier. Choose 720p or 1080p, or upgrade your plan.

From the API

Every setting above is also available to code and to AI assistants through POST /v1/generate-video. The older route POST /v1/text-to-video still works and behaves the same way.

  • Direction by id. Send a direction object of picker ids, for example { "cameraMotion": "dolly-in", "timeOfDay": "dawn" }, and Nodaro writes the matching wording into the prompt. Valid ids come from GET /v1/picker-catalogs. See Picker catalogs.
  • Frames. frameFit and frameDelivery set Frame fit and Send frames as.
  • Corrections. Corrected values come back in an adjustments array.
  • Defaults. A request without provider runs on Seedance 2 Fast. If duration is also missing, it makes a 4-second clip.
  • Character voice. On models that can speak dialogue, such as VEO 3.1, Kling 2.6, Kling 3.0, Kling 3 Omni, the Seedance 2 family and Hailuo 3, you can pass characterVoices and dialogue. The clip then speaks in the character's saved voice, for an extra 40 credits on VEO and Kling, or 25 credits per 1,000 characters on Seedance and Hailuo 3. On other models the voice is ignored, the clip still generates and no extra credits are charged.

See Run a single node and the MCP tools.

Frequently asked questions

Last updated on

On this page