> For the complete documentation index, see [llms.txt](https://docs.chatvideopro.com/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://docs.chatvideopro.com/features/image-generation/supported-image-models.md).

# Supported Image Models

Chat Video Pro includes several image models because image generation has different jobs: cinematic frames, fast drafts, product visuals, thumbnails, text-heavy designs, reference-guided edits, background removal, and upscaling.

You do not need to memorize every model. Start with what you are trying to make, then choose the model or workflow that fits the job.

{% hint style="info" %}
If you are creating a cinematic frame that may become a video, start with Studio Cinematic Lab. It wraps image generation in camera, lens, focal length, aperture, reference, aspect ratio, and model controls so you can think like a DP instead of managing raw model settings.
{% endhint %}

***

### Quick Recommendations

<table><thead><tr><th width="247">If you need...</th><th width="194">Start with...</th><th>Why</th></tr></thead><tbody><tr><td>A cinematic still or source frame for video</td><td><strong>Studio Cinematic Lab</strong></td><td>Best guided workflow for photoreal frames, references, and video-ready stills.</td></tr><tr><td>A fast, high-quality general image</td><td><strong>Nano Banana 2</strong></td><td>Strong default for most image generation and reference-guided work.</td></tr><tr><td>The hardest version of a Nano Banana prompt</td><td><strong>Nano Banana Pro</strong></td><td>Better for complex reasoning, difficult references, or harder real-world details.</td></tr><tr><td>Text inside the image</td><td><strong>GPT Image 2</strong></td><td>Best choice when signs, posters, packaging, titles, or readable words matter.</td></tr><tr><td>Highest realism or final production stills</td><td><strong>Flux 2 Max</strong></td><td>Strong detail, realism, and polished image quality.</td></tr><tr><td>Creative, affordable, flexible image generation</td><td><strong>Seedream v5</strong></td><td>Good for broad creative work, multiple aspect ratios, and efficient drafts.</td></tr><tr><td>Fast social/mobile image drafts</td><td><strong>Grok</strong></td><td>Fast image generation with unusual mobile and panoramic formats.</td></tr><tr><td>Quick low-cost previews</td><td><strong>Z-Image Turbo</strong></td><td>Good for fast concept checks and simple drafts.</td></tr><tr><td>Complex image editing with layers/masks</td><td><strong>Canvas Editor</strong></td><td>Better than plain prompting for controlled edits.</td></tr><tr><td>Transparent cutouts</td><td><strong>Background Removal</strong></td><td>Use the dedicated background-removal workflow.</td></tr><tr><td>Larger/sharper final image</td><td><strong>Image Upscaling</strong></td><td>Use a dedicated upscaler after the image is approved.</td></tr></tbody></table>

***

### The Simple Rule

Choose the model based on the hardest part of the image:

| Hardest part of the image                    | What to prioritize                                     |
| -------------------------------------------- | ------------------------------------------------------ |
| It must look like a real production frame    | Cinematic Lab, Nano Banana, or Flux.                   |
| It contains readable text                    | GPT Image 2.                                           |
| It needs a specific person, product, or look | Reference-friendly models or Cinematic Lab references. |
| It is a quick draft                          | Nano Banana 2, Grok, Seedream v5, or Z-Image Turbo.    |
| It needs precise editing                     | Canvas Editor or image-to-image mode.                  |
| It needs a transparent background            | Background Removal.                                    |
| It is already good but too small             | Image Upscaling.                                       |

The best model is not always the most expensive model. The best model is the one that solves the specific visual problem.

***

### Model Guide

#### Nano Banana 2

Nano Banana 2 is the best default for most image generation in Chat Video Pro.

Use it for:

* Cinematic images.
* Character or product concepts.
* Reference-guided image generation.
* Fast visual exploration.
* Source frames for Motion Director, Multi-Cam, AI Transitions, or Relight Scene.
* Drafting multiple directions before choosing a final look.

Choose Nano Banana 2 when you want a strong first answer quickly.

#### Nano Banana Pro

Nano Banana Pro is the version to try when the prompt is harder.

Use it for:

* Complex multi-subject scenes.
* Rare locations or real-world references.
* Images where the relationship between objects must make sense.
* Prompts that need more careful reasoning.
* Final candidates where Nano Banana 2 is close but not quite enough.

Do not use Pro for everything by default. For normal drafts, Nano Banana 2 is usually faster and more efficient.

#### GPT Image 2

GPT Image 2 is the model to try when prompt adherence, composition, or text matters.

Use it for:

* Signs, posters, book covers, labels, packaging, UI, or readable words.
* Complex compositions.
* Multi-image editing and Canvas Editor workflows.
* Weird or specific prompts where logic matters.
* Images that need careful placement of several elements.

If an image contains text that needs to be readable, switch to GPT Image 2 early.

#### Flux 2 Max

Flux 2 Max is a strong choice for high-quality, realistic, final-looking images.

Use it for:

* Final production stills.
* High-detail images.
* Photoreal texture.
* Product visuals.
* Polished image outputs where speed is less important.

Flux is a good comparison pass when a Nano Banana result is good but you want to test a more premium finish.

#### Seedream v5

Seedream v5 is a flexible creative option with broad format support.

Use it for:

* Creative concepts.
* Affordable drafts.
* Artistic or stylized images.
* Multiple aspect ratio needs.
* Text-to-image and image-editing workflows.

Seedream is a useful middle ground: more capable than a quick preview model, but still practical for iteration.

#### Grok

Grok is useful for fast image generation and mobile-first formats.

Use it for:

* Social content.
* Quick drafts.
* Mobile aspect ratios.
* Panoramic or unusual formats.
* Simple image-to-image edits.

Use Grok when speed and format flexibility matter more than maximum polish.

#### Z-Image Turbo

Z-Image Turbo is for speed.

Use it for:

* Quick visual tests.
* Low-cost iteration.
* Simple drafts.
* Early idea exploration.

Move to Nano Banana, GPT Image 2, Flux, or Seedream once the direction is worth polishing.

***

### Which Workflow Should I Use?

Model choice matters, but workflow choice comes first.

| You want to...                         | Use...                                   |
| -------------------------------------- | ---------------------------------------- |
| Create a normal image from a prompt    | Text-to-Image                            |
| Edit an existing image with a prompt   | Image-to-Image                           |
| Build a cinematic still for video work | Studio Cinematic Lab                     |
| Make a controlled layered edit         | Canvas Editor                            |
| Remove the background                  | Background Removal                       |
| Increase resolution                    | Image Upscaling                          |
| Create YouTube thumbnail concepts      | Thumbnail Mode                           |
| Animate a still image                  | Image-to-Video or Studio Motion Director |

If you are making an image that will become video, think ahead. A clean, well-framed still is easier to animate, relight, transition, or use for Multi-Cam later.

***

### Image Generation vs. Cinematic Lab

Use normal Image Generation when you want direct model control or a quick one-off image.

Use Cinematic Lab when the image is part of a production workflow.

| Use normal Image Generation when...           | Use Studio Cinematic Lab when...                                                     |
| --------------------------------------------- | ------------------------------------------------------------------------------------ |
| You need a quick image in chat.               | You are building a cinematic frame.                                                  |
| You already know which model to use.          | You want camera, lens, focal length, and aperture controls.                          |
| The image is a simple asset or draft.         | The still may become a Motion Director, Multi-Cam, AI Transition, or Relight source. |
| You want direct prompt/model experimentation. | You want references, visual consistency, and production-style look control.          |

Cinematic Lab is usually the better choice for hero frames, thumbnails, key art, style frames, and source images for Studio video workflows.

***

### Text, Logos, And Readable Words

AI image models do not all handle text equally.

Use GPT Image 2 when the image includes:

* A sign.
* A label.
* Packaging text.
* A book title.
* A poster.
* UI text.
* A thumbnail with readable words.

For text-heavy images, keep the wording short and clear. Long paragraphs inside generated images are still difficult. For final thumbnails or titles, you may get better control by generating the image background first, then adding text in Premiere, Photoshop, or another design tool.

***

### References And Image Inputs

References are useful when the output should match something: a person, product, location, wardrobe, lighting style, or art direction.

Good references are:

* Clear.
* High-resolution enough to read.
* Consistent with the desired result.
* Focused on the subject or style you want to preserve.
* Not overloaded with conflicting looks.

Use fewer references when the prompt is being ignored. Use stronger references when the model is drifting too far from the subject.

For projects with recurring characters, products, or locations, build a small reference set and reuse it. Consistency improves when the model sees the same visual anchors across generations.

***

### Speed vs. Quality

Use faster models when:

* You are brainstorming.
* You are testing composition.
* You need several visual directions.
* The image is not final.

Use higher-quality models when:

* The image is client-facing.
* The image will become a video source frame.
* Details, faces, products, or realism matter.
* You are close to the final look.

A practical workflow:

1. Draft quickly.
2. Pick the strongest direction.
3. Refine the prompt or references.
4. Generate a higher-quality final.
5. Upscale only after the image is approved.

***

### Practical Starting Points

#### If you are new

Start with:

* **Nano Banana 2** for most image generation.
* **GPT Image 2** when text or complex composition matters.
* **Cinematic Lab** for cinematic frames and video-ready stills.
* **Z-Image Turbo** for quick drafts.

#### If you are making video source frames

Start with:

* Cinematic Lab for controlled camera/lens look.
* Nano Banana 2 for fast photoreal exploration.
* GPT Image 2 if the frame contains readable text.
* Flux 2 Max when you want to compare a more polished realism pass.

Then send the result into Motion Director, Multi-Cam, AI Transitions, or Relight Scene.

#### If you are making thumbnails

Start with:

* Thumbnail Mode for thumbnail-specific composition.
* GPT Image 2 when text or layout precision matters.
* Nano Banana or Flux for strong faces, scenes, or image backgrounds.

For final thumbnail text, consider adding the text manually after generation for maximum control.

#### If you are editing an existing image

Start with:

* Image-to-Image for prompt-based edits.
* Canvas Editor for controlled edits with masks, layers, or annotations.
* Background Removal for cutouts.
* Image Upscaling after the edit is approved.

***

### Troubleshooting

#### The image looks generic

Add more specific visual direction: subject details, setting, lighting, material, texture, mood, and composition. If it should feel like a real frame, try Cinematic Lab.

#### Text in the image is wrong

Use GPT Image 2 or add the final text manually in a design tool. Keep generated text short.

#### The model ignores my reference

Use a clearer reference, reduce conflicting references, or make sure the prompt does not contradict the image. For recurring subjects, reuse the same reference set across generations.

#### The model follows the reference too much

Use fewer references or write a stronger prompt describing what should change. If all references look similar, the model may treat them as strict instructions.

#### The image is too small or soft

Do not regenerate endlessly just for resolution. Once the image is approved, use Image Upscaling.

#### I am not sure which model to pick

Start with Nano Banana 2. If the problem is text, switch to GPT Image 2. If the problem is final realism, compare Flux 2 Max. If the problem is speed, use Z-Image Turbo or Grok. If the problem is cinematic framing, use Cinematic Lab.

***

### Next Steps

* Use [Text-to-Image](/features/image-generation/text-to-image.md) for normal prompt-based image generation.
* Use [Image-to-Image](/features/image-generation/image-to-image.md) when editing an existing image.
* Use Studio [Cinematic Lab](/features/studio/cinematic-lab.md) for cinematic stills and video-ready frames.
* Use [Canvas Editor](/features/image-generation/canvas-editor.md) for controlled edits.
* Use [Image Upscaling](/features/image-generation/image-upscaling.md) after the image is approved.
