Supported Video Models
Chat Video Pro supports multiple state-of-the-art AI video generation models, each optimized for different use cases.

Chat Video Pro includes several video models because no single model is best at everything. Some are better for dialogue. Some are better for cinematic camera movement. Some are better for fast drafts, longer clips, reference consistency, or mobile-first content.
This page is a practical chooser guide. You do not need to memorize every model. Start from the type of shot you want, then choose the model that fits the job.
If you are using Studio, the workflow often chooses the model path for you. For example, Motion Director uses Kling 3.0 Pro Image-to-Video, AI Transitions uses Kling O3 transition models, and Add Effects uses Kling O3 VFX. Use this page when you are choosing models directly from the Video Generation model selector.
Quick Recommendations
Dialogue, speech, or generated audio
Veo 3.1 or Veo 3.1 Fast
Strongest choice when audio and lip-sync matter.
Fast lower-cost Veo drafts without audio
Veo 3.1 Lite
Good for visual drafts, B-roll ideas, and transitions when you will add sound in Premiere.
Cinematic motion and strong general quality
Kling 3.0 Pro or Kling O3 Pro
Strong camera movement, action, and polished motion.
Longer cinematic clips without audio
Sora 2 or Sora 2 Pro
Good for cinematic B-roll and longer visual generations.
Natural motion with ambient audio
Seedance 2
Good all-around model for natural motion and audio-enabled scenes.
Action, sports, or fast movement
Hailuo 03
Strong for dynamic action and energetic movement, at 2K with native audio.
Reference from a video or audio clip
Hailuo 03 Reference or Seedance 2.5 Reference
Both take reference video and reference audio, not just stills. Hailuo 03 caps at 12 files combined and renders 2K; Seedance 2.5 accepts up to 50 references combined and runs to 30s at 720p.
Fast content with audio built in
Grok Imagine 1.5
Fast generations up to 1080p with audio always on.
High-resolution 1080p-style outputs
Wan 2.7
Good when clarity, flexible aspect ratios, or reference workflows matter.
A guided camera move from a still image
Studio Motion Director
Easier than hand-prompting image-to-video movement.
A polished transition between two frames
Studio AI Transitions
Easier than manually choosing a transition model.
Native 4K video output
Seedance 2 (4K resolution pill)
Full 3840×2160 without separate upscale.
Fast Seedance iterations
Seedance 2 Mini
Seedance-quality at lower cost.
Audio and video in one pass
Google Omni Flash
Synchronized audio.
Fast mobile-first with lipsync
Grok Imagine 1.5
1080p image-to-video with lipsync.
2K output with native audio
Hailuo 03
2K on every tier, 5–15s, stereo audio always on.
A clip of 15–20s at 1080p
Flux 3
5–20s with native audio at 720p or 1080p.
A clip that runs past 20s
Seedance 2.5
4–30s in a single pass with native audio, at 480p or 720p.
A reference set larger than 12 files
Seedance 2.5 Reference
Accepts up to 50 references combined across images, videos, and audio.
Ultra-wide 2:1 or 21:9 framing
Flux 3
Eight aspect ratios, and no other video model offers 2:1.
A silent cinematic plate with a cheap draft tier
Luma Ray 3.2
5s or 10s, 540p/720p/1080p, no audio pill anywhere — score it in Premiere.
Restyle a clip with control over how far it diverges
Luma Ray 3.2 Edit
Video-to-video with a 4-way divergence dial: Default, Adhere, Balanced, Reimagine.
The Simple Rule
Choose the model based on the hardest part of the shot:
A person speaking
Audio and lip-sync.
Complex camera movement
Motion quality and scene understanding.
A client-ready cinematic insert
Quality and consistency.
A quick concept
Speed and cost.
A character must look the same
Reference mode or Studio workflow.
Two frames need to connect
Transition mode or Studio AI Transitions.
The clip must be vertical/mobile
Aspect ratio support.
The best model is not always the highest-quality model. The best model is the one that solves the specific problem in the shot.
Model Guide
Veo 3.1
Use Veo 3.1 when the shot needs audio, dialogue, or speaking characters.
Best for:
Talking-head concepts.
Product explainers with speech.
Scenes where sound matters.
Short dialogue tests.
Image-to-video or transition shots where audio should be part of the generation.
Use Veo 3.1 Fast when you want quicker iterations. Use the regular Veo 3.1 path when quality matters more than speed.
Veo 3.1 Lite
Use Veo 3.1 Lite when you want a lower-cost Veo-style visual draft and do not need generated audio.
Best for:
Silent B-roll concepts.
Visual drafts before adding voiceover or music in Premiere.
Budget-conscious text-to-video, image-to-video, or transition tests.
Shots where 720p or 1080p is enough.
Avoid Lite when the prompt depends on spoken dialogue or synchronized audio. Add sound in Premiere instead.
Kling 3.0
Use Kling 3.0 when you want strong cinematic movement, flexible shot types, and polished video generation.
Best for:
Camera moves.
Dynamic product shots.
Cinematic B-roll.
Image-to-video with strong motion.
Shots that need native audio but are less dialogue-focused than Veo.
Kling 3.0 is a strong general-purpose choice for classic text-to-video and image-to-video generation.
Kling O3
Use Kling O3 when motion quality, scene understanding, or transition quality is the priority.
Best for:
Advanced motion.
High-quality transitions.
Reference-driven shots.
VFX-oriented generations.
Complex scenes where the model needs to preserve structure.
If you are creating transitions, consider Studio AI Transitions instead of manually selecting an O3 transition model. The Studio workflow gives you transition styles and better prompting structure.
Sora 2
Use Sora 2 when you want cinematic visual quality and do not need generated audio.
Best for:
Establishing shots.
Atmospheric B-roll.
Longer visual clips.
Cinematic concepts where dialogue is not required.
Choose Sora 2 Pro when you want the higher-quality Sora option and the extra cost makes sense.
Seedance 2
Use Seedance 2 when you want natural motion with native audio and a balanced all-around video model. Select the 4K resolution pill for full 3840×2160 output without a separate upscale pass.
Best for:
Natural movement.
Ambient audio scenes.
Cinematic clips with sound.
Reference mode when you need subject consistency.
Wide or cinematic aspect ratios.
Native 4K delivery when the resolution pill is enabled.
Use Seedance 2 Mini for faster, lower-cost iterations when you want Seedance-quality motion without the full standard price.
Use Seedance 2 Fast when you want the same general family with quicker turnaround.
Seedance 2.5
Use Seedance 2.5 when the clip needs to run longer than the rest of the Seedance family, or when a reference set is bigger than other models will take.
Seedance 2.5 ships alongside Seedance 2, not as a replacement. It runs to 30 seconds but stops at 720p, where Seedance 2's standard tier reaches 1080p and 4K — so which one is "better" depends entirely on whether you need running time or resolution. Both sit in the same Seedance submenu.
Best for:
Clips of 4–30 seconds in a single pass, with native audio.
Letting the model choose the length — Auto is the default duration.
Several shots inside one generation, sequenced by describing the cuts in the prompt.
Start-frame-to-end-frame morphs that need more than 15 seconds.
Reference-driven work with a large or mixed reference set.
There are two picker entries: Seedance 2.5 (text-to-video and image-to-video on one entry) and Seedance 2.5 Reference.
Duration: 4–30 seconds, whole seconds, or Auto — and Auto is the default.
Resolution: 480p or 720p only, defaulting to 720p. There is no 1080p and no 4K on 2.5.
Audio: native, generated with the picture, On by default and genuinely toggleable — the same pill the Seedance 2 tiers use. It is ambient and scene audio, not dialogue-grade lip-sync.
Aspect ratios: auto (default), 21:9, 16:9, 4:3, 1:1, 3:4, 9:16 — the same seven the Seedance 2 tiers offer.
Two images: the second becomes the end frame on the same image-to-video endpoint. Nothing re-routes, and there is no separate transition entry.
References: the composer shows 9 image, 3 video, and 3 audio slots. The model itself accepts up to 50 references combined across all three types, so that budget is headroom at the API rather than something today's composer can fill.
There is no Seedance 2.5 Fast and no Seedance 2.5 Mini — 2.5 has one quality tier plus Reference. Those speed variants exist only for Seedance 2. Multi-shot sequencing is likewise prompt-driven: describe the cuts you want. There is no multi-shot control anywhere in the UI.
Seedance 2.5 is billed by tokens rather than a flat per-second rate: $0.0214 per 1,000 tokens, where tokens are height × width × seconds × 24 ÷ 1024. Because the frame dimensions are in the formula, the per-second cost varies with the aspect ratio you choose. At 16:9 that works out to roughly $0.4730 per second at 720p and $0.2205 at 480p, so a 30-second 720p clip is about $14. On the Reference tier, attaching a video reference applies a ×0.6 multiplier — but the input video's duration is billed alongside the output's, so a long source clip is not free.
Google Omni Flash
Find Omni Flash under Generate Media → Video → Google → Omni Flash.
Use Google Omni Flash when you want synchronized audio and video generated in one pass.
Best for:
Text-to-video and image-to-video with native audio.
720p clips from 3–10 seconds.
16:9 or 9:16 social and landscape formats.
Scenes where dialogue, ambient sound, or lip-sync should match the visuals.
Omni Flash Reference accepts 2–7 images. When you attach multiple images, the app auto-selects Reference mode.
Omni Flash Edit is a video-to-video path under Google alongside Veo Extend. Use it when you want to edit or extend an existing clip while keeping Omni Flash's synchronized audio behavior.
Hailuo 03
Use Hailuo 03 when action and motion are the main challenge, or when you want 2K with sound already on it.
Best for:
Sports.
Fast movement.
Stunts.
Energetic social clips.
Dynamic camera action.
Every Hailuo 03 tier renders at 2K — the only resolution offered — for 5 to 15 seconds, with native stereo audio always on. There is no quality tier to choose and no audio toggle to forget.
Hailuo 03 (text-to-video): aspect ratios 21:9, 16:9, 4:3, 1:1, 3:4, 9:16, default 16:9. Prompts up to 2000 characters.
Hailuo 03 (image-to-video): attach one image to animate it, or two to interpolate a start frame into an end frame. There is no aspect control here — the output follows the source image, so crop the still first if you need a different shape.
Hailuo 03 Reference: up to 9 reference images, 3 reference video clips (2–15s each, 50 MB total), and 3 reference audio clips (2–15s each, 15s combined), to a maximum of 12 files in total. Audio can never be the only reference. Aspect ratio defaults to Adaptive, which follows your references, or pick a concrete ratio.
It is less of a first choice for dialogue-heavy work — reach for Veo 3.1 or Grok Imagine 1.5 when lip-sync matters. Hailuo 03 also takes its time: it renders 2K with audio and can sit in the queue for a while, so send it and carry on editing rather than waiting on it.
Wan 2.7
Use Wan 2.7 when you want flexible aspect ratios, high-resolution visual output, or reference-driven generation without generated audio.
Best for:
Clean visual generations.
Flexible formats.
Reference mode with more image inputs.
Start/end interpolation.
Instruction-style video editing tests.
Plan to add sound later if the final edit needs audio.
Flux 3
Use Flux 3 when you want a long shot at 1080p that already has sound on it.
Best for:
Clips of 5–20 seconds at 720p or 1080p.
Long shots that need native audio in the same pass.
Ultra-wide framing: Flux 3 offers eight aspect ratios, and no other video model in the app offers 2:1.
Start-frame-to-end-frame transitions that need room to play out.
1080p delivery without thinking about quality tiers.
Flux 3 comes from Black Forest Labs and appears in the picker under the Flux provider group as simply Flux 3.
Duration: 5–20 seconds, whole seconds only, default 8s. Most other families stop at 15 seconds; Seedance 2.5 runs further, to 30s, but caps at 720p.
Resolution: 720p or 1080p, defaulting to 1080p. There is no 4K tier and no separate quality toggle — resolution is the quality setting.
Audio: native and generated with the picture. It ships On and, unlike Hailuo 03 or Grok, it can genuinely be switched off. Black Forest Labs recommends describing the ambient sound in your prompt.
Aspect ratios: auto, 21:9, 2:1, 16:9, 4:3, 1:1, 3:4, 9:16. Text-to-video defaults to 16:9; image-to-video defaults to auto but is not locked to the source image.
Two images: attach exactly two and Flux 3 routes to a dedicated first/last-frame endpoint, treating the first as the start frame and the second as the end frame.
Flux 3 has no reference tier. Attach one image and only the image-to-video tier is offered — the text-to-video tier hides. Attach three or more images and Flux 3 leaves the picker entirely, handing off to a reference model such as Veo 3.1 Reference. There is also no seed input, no negative prompt, and no 4K.
Cost scales purely with duration: $0.17 per second at 720p and $0.29 per second at 1080p, the same across all three Flux 3 endpoints. A 20-second 1080p clip is $5.80, so draft short at 720p and stretch the duration only on the take you are keeping.
Luma Ray 3.2
Use Luma Ray 3.2 when you want a silent cinematic clip you will score yourself, a cheap drafting tier, or an edit pass over existing footage with explicit control over how far it strays.
Best for:
Silent plates and B-roll you plan to sound-design in Premiere.
Cheap drafting — the 540p tier prices well below the 720p default.
A quick, inexpensive start/end interpolation from two stills.
Restyling or editing an existing clip with Luma Ray 3.2 Edit and its divergence dial.
The family appears in the picker under the Luma provider group. The generation tiers are labelled Luma Ray 3.2; the edit model is Luma Ray 3.2 Edit in the video-to-video section, and it is also selectable inside Studio → Add Effects.
Duration: 5s or 10s on text-to-video and Edit, default 5s. Image-to-video is 5s only — 10s needs multi-keyframe input that Chat Video Pro does not expose.
Resolution: 540p, 720p, or 1080p on every tier, defaulting to 720p. Resolution is the only quality setting.
Audio: none, on any tier. There is no audio pill anywhere in the family — the models have no audio parameters, so nothing can switch sound on. Build the soundtrack in Premiere.
Aspect ratios: 16:9 (default), 9:16, 1:1, 4:3, 3:4, 21:9 — no auto. The Edit model has no aspect setting at all; its output follows the source clip.
Two images: attach exactly two and the second becomes the end frame on the same image-to-video endpoint — unlike Flux 3, nothing re-routes to a separate first/last-frame endpoint.
The Edit divergence dial: Default (Luma decides), Adhere (closest to the source), Balanced, Reimagine (diverges the most).
Luma Ray 3.2 has no reference tier and narrows the same strict way Flux 3 does: attach one image and only the image-to-video tier is offered; attach three or more images (or use Element tags) and the family leaves the picker entirely. There is also no seed, no negative prompt, and nothing above 1080p. Do not confuse it with Studio Reframe — that is a separate Luma-powered Studio surface with its own page, and it never appears in the chat video-to-video picker.
Pricing is per second, constant across 5s and 10s, and differs per endpoint: text-to-video $0.20/s at 720p ($0.10 at 540p, $0.40 at 1080p), image-to-video $0.06/s at 720p ($0.03 at 540p, $0.24 at 1080p), Edit $0.216/s at 720p ($0.144 at 540p, $0.432 at 1080p). A 10s text-to-video clip at the default tier is $2.00; a 10s Edit pass is $2.16.
Grok Imagine 1.5
Use Grok Imagine 1.5 when speed matters and you want the clip to come back with sound already on it.
Best for:
Up to 1080p on both text-to-video and image-to-video.
Lipsync on single-image animation.
Fast drafts and social clips.
Quick content experiments.
Audio is always on across the Grok video tiers — there is no toggle. Duration runs 1 to 15 seconds.
Grok Reference accepts up to 7 images for reference-driven generation, at 480p or 720p.
Grok Imagine 1.5 is useful for fast exploration. For final cinematic polish, compare against Veo, Kling, Sora, or Seedance.
Which Mode Should I Use?
Model choice matters, but mode choice comes first.
Only a prompt
Text-to-Video
One still image
Image-to-Video
A start frame and end frame
Transition Mode or Studio AI Transitions
Several images of the same subject
Reference Mode
An existing video that should continue
Generative Extend
A still image that needs directed camera movement
Studio Motion Director
A video that needs cleanup, VFX, relight, reshoot, or upscale
Studio
If you are not sure, start with the inputs. The number and type of files you attach usually determines the best mode.
Audio Guide
Dialogue or lip-sync
Veo 3.1, Veo 3.1 Fast, Google Omni Flash
I2V lipsync from a still
Grok Imagine 1.5
Ambient sound or general audio
Kling 3.0, Kling O3, Seedance 2, Seedance 2.5, Google Omni Flash, Hailuo 03, Flux 3
Audio on a clip longer than 15s
Flux 3 (to 20s, up to 1080p) or Seedance 2.5 (to 30s, up to 720p)
Silent visual draft
Veo 3.1 Lite, Sora 2, Wan 2.7, Seedance 2 Mini, Luma Ray 3.2
Final sound design
Generate visuals first, then finish audio in Premiere
Audio generation can be useful, but it is not always the best final audio. For client work, you may still want to add dialogue, voiceover, music, or sound effects in Premiere.
Pro vs. Fast vs. Standard
Some model families have quality or speed variants.
Use the higher-quality option when:
The shot is for delivery.
The prompt is complex.
Character or product consistency matters.
You are generating a final hero shot.
Use the faster or lighter option when:
You are exploring ideas.
You are testing prompts.
You expect to regenerate several times.
You care more about speed or cost than final polish.
A good workflow is to draft with a faster model, then regenerate the best prompt with a higher-quality model.
When To Use Studio Instead
Studio is better than manual model selection when the task already has a clear workflow.
Create a cinematic still before video work
Cinematic Lab
Animate a still with a camera move
Motion Director
Create a polished transition
AI Transitions
Generate alternate angles
Multi-Cam
Transfer motion to a character image
Motion Capture
Add VFX to a video
Add Effects
Change lighting
Relight Scene
Reshoot a short segment
Reshoot
Finish resolution
Upscale
Studio does more of the prompt and model setup for you. Direct model selection gives you more control when you already know exactly which generation mode you want.
Practical Starting Points
If you are new
Start with one of these:
Veo 3.1 Fast for quick audio-enabled video tests.
Kling 3.0 Pro for cinematic motion and image-to-video.
Veo 3.1 Lite for lower-cost silent drafts.
Studio Motion Director if you already have a strong still image.
If you are making B-roll
Try:
Sora 2 for cinematic visual clips.
Kling 3.0 or Kling O3 for stronger movement.
Veo 3.1 Lite for lower-cost drafts.
Seedance 2 if audio/ambient sound is useful.
If you are making social content
Try:
Grok Imagine 1.5 for fast turnaround, up to 1080p, and I2V lipsync.
Hailuo 03 for action and energy at 2K with sound.
Kling for polished camera movement.
Veo 3.1 when dialogue or sound matters.
If you are making client-facing shots
Use faster models to explore, then move the best idea into a quality pass:
Draft with a faster or lighter model.
Refine the prompt.
Regenerate with a higher-quality model.
Use Studio tools for cleanup, transitions, relight, or upscale.
Troubleshooting
The model I expected is not available
The model selector changes based on what you attach. No attachments show text-to-video models. One image shows image-to-video options. Two images may activate transition mode. Multiple reference images may show reference models.
The result has no audio
Check whether the selected model supports audio. Veo 3.1, Kling, Seedance (2 and 2.5 alike), Google Omni Flash, Hailuo 03, and Flux 3 can generate audio — on Hailuo 03 it is always on, with no toggle. On Flux 3 and on the Seedance tiers the audio pill defaults to On but can be switched off, and Chat Video Pro remembers that choice across models, so if you turned audio off on an earlier generation it will still be off here. Grok Imagine 1.5 supports lipsync on image-to-video. Sora 2, Veo 3.1 Lite, and Wan 2.7 are better treated as visual models unless the app shows an audio option for your selected mode. Luma Ray 3.2 is silent on every tier — there is no audio pill in the family at all, so a silent result there is expected, not a bug.
The model looks wrong for my use case
Switch based on the failure:
Weak dialogue or lip-sync
Veo 3.1
Weak motion
Kling, Seedance, or Hailuo 03
Clip is too short
Flux 3 (to 20s at 1080p) or Seedance 2.5 (to 30s at 720p)
Need faster drafts
Veo 3.1 Lite, Seedance 2 Mini, Grok Imagine 1.5, or a Fast/Standard variant
Need stronger transition
Studio AI Transitions
Need more controlled image animation
Studio Motion Director
Need a cleaner final
Upscale after the creative result is approved
The model list changes over time
Chat Video Pro adds and updates models as providers improve. If the in-app selector differs from this page, trust the app. This guide is meant to help you choose the right kind of model, not memorize every technical option.
Next Steps
Use Text-to-Video when starting from a prompt.
Use Image-to-Video when animating one still.
Use Transition Mode or AI Transitions when connecting two frames.
Use Reference Mode when subject consistency matters.
Use Studio for guided production and post-production workflows.
Last updated