For the complete documentation index, see llms.txt. This page is also available as Markdown.

Text-to-Image

Create images from scratch using text descriptions. Simply describe what you want to see, and Chat Video Pro generates it using AI.

Text-to-Image creates a still image from a written prompt. Use it when you want a new image, concept, thumbnail idea, product visual, style frame, background, or reference asset and you do not already have a source image to edit.

Chat Video Pro gives you two normal Text-to-Image paths: a fast chat path for quick ideas, and a full-control Generate Media path when you need model, aspect ratio, resolution, and quality settings. For cinematic stills and video-ready source frames, use Studio Cinematic Lab.


When To Use Text-to-Image

Use Text-to-Image when you want to create a still from scratch:

  • A quick concept image.

  • A product or brand visual.

  • A thumbnail background.

  • A mood board image.

  • A social post image.

  • A background plate.

  • A first draft before image editing.

  • A source image for later video generation.

Choose another workflow when:

You want to...
Use instead

Create a cinematic frame for video work

Studio Cinematic Lab

Edit an existing image

Image-to-Image

Make a controlled layered edit

Canvas Editor

Remove a background

Background Removal

Increase resolution

Image Upscaling

Build a YouTube thumbnail

Thumbnail Mode

Animate the image into video

Image-to-Video or Studio Motion Director


Quick Path vs. Full Control

Quick Path

Use the quick path when you want a fast image during a conversation.

  1. Set your default image model in the top-right model selector.

  2. Type a natural language request in chat.

  3. Example: Generate an image of a cozy cabin in the snow at night.

  4. Chat Video Pro generates the image using default settings.

Best for:

  • Fast ideas.

  • Casual requests.

  • Conversation flow.

  • Images where exact aspect ratio and resolution are not critical.

Limitations:

  • Less control.

  • Uses default settings.

  • Not ideal for production-specific format requirements.

Full Control

Use Generate Media mode when the image needs specific settings.

  1. Enable Generate Media in the composer.

  2. Choose an image model.

  3. Choose aspect ratio and resolution/quality.

  4. Enter a more detailed prompt.

  5. Generate and review.

Best for:

  • Production work.

  • Specific aspect ratios.

  • Higher-quality outputs.

  • Model comparisons.

  • Images that may become thumbnails, key art, or video source frames.


Text-to-Image vs. Cinematic Lab

Use normal Text-to-Image when you want direct model control or a quick image.

Use Cinematic Lab when the still needs to feel like a real production frame.

Use Text-to-Image when...
Use Studio Cinematic Lab when...

You need a quick image or concept.

You need a cinematic still with camera/lens control.

You know which image model you want.

You want the workflow to guide the visual look.

The image is a draft or simple asset.

The image may become a Motion Director, Multi-Cam, AI Transition, or Relight source.

You want to manually write the whole image prompt.

You want camera body, lens, focal length, aperture, references, and model controls.

The practical rule: use Text-to-Image for general image generation, and use Cinematic Lab when the still is part of a video or production workflow.


What A Good Prompt Includes

A strong image prompt usually describes:

Prompt element
What to include

Subject

The main person, object, place, or idea.

Setting

Where the image takes place.

Composition

Close-up, wide shot, centered product, low angle, overhead, etc.

Lighting

Golden hour, studio lighting, neon, soft window light, moody shadows.

Style

Photoreal, editorial, cinematic, product photo, illustration, graphic design.

Details

Materials, textures, colors, wardrobe, props, background elements.

Mood

Calm, energetic, premium, mysterious, cozy, futuristic.

Useful structure:

You do not need to include everything every time. Add the details that actually matter for the image.


Prompt Examples

Cinematic Concept Image

Why it works:

  • Clear subject and setting.

  • Lighting and mood are specific.

  • Composition gives the model a shot shape.

Product Visual

Why it works:

  • Product, material, surface, lighting, and style are all clear.

  • The image has a usable commercial direction.

Social Background

Why it works:

  • It tells the model the final use.

  • It reserves space for text.

  • It avoids overloading the image with detail.

Thumbnail Concept

Why it works:

  • It includes layout.

  • It leaves room for text.

  • It uses thumbnail-specific visual language.


Weak Prompts To Avoid

Too vague:

Better:

Missing composition:

Better:

No use case:

Better:


Choosing A Model

Use the model based on the hardest part of the image.

Need
Good starting point

Strong general image generation

Nano Banana 2

Harder prompt or reference reasoning

Nano Banana Pro

Readable text in the image

GPT Image 2 or Ideogram V4 Fast

Final realism and detail

Flux 2 Max

Typography, posters, UI mockups, region-precise edits

Seedream 5.0 Pro

Fast mobile/social formats

Grok

Quick design drafts with readable text

Ideogram V4 Fast

Cinematic frame for video

Studio Cinematic Lab

For a deeper chooser guide, see Supported Image Models.


Aspect Ratio Guide

Choose aspect ratio based on where the image will be used.

Aspect ratio
Best for

16:9

Video frames, YouTube thumbnails, website headers, landscape images.

9:16

Shorts, Reels, TikTok, vertical stories, phone screens.

1:1

Square social posts, profile-style images, balanced compositions.

4:5

Instagram feed portraits and social graphics.

21:9

Cinematic widescreen frames and banners.

4:3 or 3:4

Editorial, vintage, portrait, or alternate framing.

Pick the final deliverable shape before generating. Cropping after generation can cut off important subjects, text areas, or composition lines.


Best Practices

Describe The Image You Need, Not Just The Topic

"A watch" gives the model a topic. "Premium product photo of a black watch on a dark reflective surface with rim light" gives it an image.

Include Composition Early

Composition controls whether the image is usable. Say:

  • Close-up portrait.

  • Wide establishing image.

  • Centered product shot.

  • Low-angle hero shot.

  • Overhead flat lay.

  • Empty space on the right for text.

Save Text For The Right Workflow

If the image needs readable words, use GPT Image 2 or add the final text manually after generation. For thumbnails and graphics, generating the background first and adding final text yourself often gives the cleanest result.

Use References For Consistency

If a character, product, or brand look must stay consistent, attach reference images in a workflow that supports them or use Cinematic Lab references.

Generate A Few Directions Before Polishing

Do not over-optimize the first result. Generate a few directions, pick what is working, then refine.

Upscale Last

Do not use upscaling to fix a bad image. Choose the image first, then use Image Upscaling if it needs more resolution.


Example Workflows

Quick Concept Image

  1. Use the quick path in chat.

  2. Ask for a simple visual idea.

  3. If the direction works, regenerate with full control or move into Cinematic Lab.

Production Still

  1. Use Generate Media or Cinematic Lab.

  2. Choose the final aspect ratio.

  3. Write a detailed prompt with composition and lighting.

  4. Generate several directions.

  5. Use the best image as a source for Motion Director, Multi-Cam, or AI Transitions.

Thumbnail Background

  1. Choose 16:9.

  2. Prompt for strong contrast and empty title space.

  3. Avoid asking the model to create final text unless using GPT Image 2.

  4. Add final thumbnail text manually for control.

Product Visual

  1. Describe the product, material, surface, and lighting.

  2. Use a clean composition.

  3. Keep the prompt focused on the product.

  4. Use Image-to-Image or Canvas Editor for refinements.


Troubleshooting

The image feels generic

Add specific composition, lighting, material, and mood. Avoid one-word prompts. If you want a production frame, try Cinematic Lab.

The image has bad text

Use GPT Image 2 or add the text manually after generation. Keep generated text short.

The image is the wrong shape

Set the aspect ratio before generating. If the final platform is vertical, generate vertical from the start.

The model ignored an important detail

Move the detail earlier in the prompt and remove competing instructions. If the detail is a person, product, logo, or brand look, use references.

The output looks too AI-generated

Add concrete physical details: material, texture, imperfect surfaces, realistic lighting, lens feel, and environment. Cinematic Lab can help if you want a more grounded production-frame look.

The result is close but not final

Use the right follow-up workflow:

Problem
Better next step

Need to edit part of the image

Image-to-Image or Canvas Editor

Need a transparent cutout

Background Removal

Need higher resolution

Image Upscaling

Need motion

Image-to-Video or Motion Director

Need a different angle

Multi-Cam



Next: If the still needs to become a video shot, use Image-to-Video or Studio Motion Director.

Last updated