For the complete documentation index, see llms.txt. This page is also available as Markdown.

Reference Mode

Generate videos with consistent characters using 1-9 reference images. Perfect for maintaining character appearance across multiple shots, creating talking head videos, or ensuring visual consistency

Reference Mode generates video using one or more images as visual anchors. Instead of asking the model to invent a character, product, outfit, location, or style from text alone, you provide images that show what should stay consistent.

Use it when identity matters: a recurring character, presenter, product, mascot, brand object, wardrobe, location, or visual look.

If you attach exactly two images and the app switches to Transition Mode, that means Chat Video Pro is treating them as a start frame and end frame. To use the images as references instead, choose a reference-capable model such as Kling O3 Reference, Seedance 2 Reference, Seedance 2.5 Reference, Wan 2.7 Reference, Veo 3.1 Reference, Google Omni Flash Reference (2–7 images), Grok Reference (up to 7 images), or Hailuo 03 Reference (up to 9 images, plus reference video and audio clips).


When To Use Reference Mode

Use Reference Mode when you want a generated video to follow visual examples:

  • A presenter should look like the same person across shots.

  • A product should keep its shape, color, and design.

  • A character should remain recognizable.

  • A location, wardrobe, or brand style should carry across generations.

  • You need multiple videos that feel like the same campaign.

  • Text-only prompting is not enough to preserve the subject.

Choose another workflow when:

You have...
Use instead

One image that should become the first frame of a video

Image-to-Video or Studio Motion Director

Two images that should connect as start/end frames

Transition Mode or Studio AI Transitions

A still image that needs alternate camera angles

Studio Multi-Cam

A cinematic still that needs to be created first

Studio Cinematic Lab

Existing video to edit

Studio


How It Works

  1. Enable Generate Media.

  2. Attach reference images of the same subject, product, character, or look.

  3. Choose a reference-capable model.

  4. Write a prompt describing the scene, action, camera, and mood.

  5. Configure duration, aspect ratio, resolution, and audio options when available.

  6. Generate the video.

  7. Review whether the subject stayed consistent.

The references provide identity and visual direction. The prompt provides the new scene and action.


Reference Mode vs. Transition Mode

This is the most common point of confusion.

If the images are...
Use...

The first and last frame of a shot

Transition Mode

Examples of the same person/product/character

Reference Mode

Two frames you want to connect with a polished style

Studio AI Transitions

A single still you want to animate

Image-to-Video or Motion Director

Transition Mode asks: How should image A become image B?

Reference Mode asks: What should stay consistent while the model creates a new shot?


Choose Strong Reference Images

Good references are clear, consistent, and useful.

Use images that show:

  • The same person, product, or character.

  • A clear face, silhouette, product shape, or key design detail.

  • Different useful angles when possible.

  • Similar identity even if pose, expression, or lighting changes.

  • Enough resolution for the model to read details.

  • The most important visual traits you want preserved.

Avoid references that are:

  • Blurry, dark, or low quality.

  • Different people or different products.

  • Contradictory styles.

  • Extreme angles only.

  • Heavily filtered or distorted.

  • Full of unrelated background clutter.

  • Too many images that fight each other.

More references are not always better. A small set of clean references often works better than a large set of mixed-quality images.


How Many References To Use

Use the smallest set that explains the subject.

Reference count
Best for

1 image

Simple product, logo-like subject, or one clear character anchor.

2-3 images

Most people, products, presenters, and brand subjects.

4-7 images

Complex characters, varied angles, or stronger identity preservation.

8-9 images

When the selected model supports it and you have genuinely useful angle/style variety.

If results ignore your references, try better references before adding more. If results copy the references too literally, reduce the count or use more varied images.


What To Prompt

Do not spend the whole prompt describing the character if the reference images already show them. Use the prompt to describe the new shot.

Focus on:

Prompt layer
What to describe

Setting

Where the subject is now.

Action

What the subject does.

Camera

Shot size, movement, and angle.

Mood/style

Cinematic, commercial, documentary, playful, dramatic.

Audio/dialogue

Only if the selected model supports generated audio.

Preservation

What must stay consistent: face, outfit, product shape, logo, color.

Useful structure:


Prompt Examples

Presenter Video

Character Scene

Product Clip

Brand Mascot


Weak Prompts To Avoid

Only describes the reference:

Better:

Too vague:

Better:

Conflicts with the reference:

Better:

If you want to change the subject itself, use image editing or another workflow. Reference Mode is strongest when references are meant to remain recognizable.


Choosing A Model

Choose based on what matters most.

Need
Good starting point

Dialogue or talking head with references

Veo 3.1 Reference

Flexible duration and strong reference quality

Kling O3 Reference

More reference images and no generated audio

Wan 2.7 Reference

Reference video with native audio/ambient sound

Seedance 2 Reference

A reference clip that needs to run past 15 seconds

Seedance 2.5 Reference — 4–30s (or Auto) at 480p or 720p, with native audio

A reference set larger than 12 files

Seedance 2.5 Reference — the model accepts up to 50 references combined across images, videos, and audio, with no per-type caps

2–7 references with synchronized audio

Google Omni Flash Reference

Up to 7 references for fast I2V with audio

Grok Reference

Reference material that is a video clip or an audio clip, not a still

Hailuo 03 Reference — up to 9 images, 3 video clips (2–15s each, 50 MB total), and 3 audio clips (2–15s each, 15s combined), 12 files in total. Audio can never be the only reference

Guided character/product still creation before video

Studio Cinematic Lab

For a broader model chooser, see Supported Video Models.

Seedance 2.5 Reference — two numbers, both true. The composer shows 9 image slots, 3 video slots, and 3 audio slots, the same layout every Seedance reference tier uses. The model itself accepts up to 50 references combined across all three types, with no per-type caps. That larger budget is headroom at the API rather than something the current composer can fill, so plan around the 9 / 3 / 3 slots you can actually see.

Two rules to keep in mind. Audio can never be the only reference — attach at least one image or video alongside it. And a video reference changes the bill: the per-second rate drops by a ×0.6 multiplier, but the input video's duration is billed as well as the output's, so a long source clip is not free. Cite references in the prompt as [Image1], [Video1], [Audio1].

Seedance 2 Reference is the tier to use when you need above 720p: 2.5 stops at 720p, and the Seedance 2 standard reference tier does not.


When To Use Studio Instead

Studio is often better when the task has a more specific creative shape.

Goal
Better Studio workflow

Create a consistent cinematic still first

Cinematic Lab

Animate a still with a camera move

Motion Director

Create alternate angles of a subject

Multi-Cam

Transition between two intentional frames

AI Transitions

Transfer motion to a character image

Motion Capture

Use Reference Mode when you want direct model control from the composer. Use Studio when you want the workflow to guide the prompt, model, and asset setup.


Best Practices

Build A Small Reference Set

For recurring people, products, or characters, keep 3-5 strong images in your Library or Recents. Reuse the same set across generations for more consistent results.

Keep References Focused

Do not mix unrelated styles unless style mixing is the goal. A product render, a blurry phone photo, and a stylized illustration may confuse the model if they are all meant to define the same subject.

Prompt The Scene, Not The Biography

References handle appearance. Your prompt should direct the new shot: where the subject is, what they are doing, how the camera moves, and what the mood is.

Preserve What Matters

If a detail must stay consistent, name it:

Expect Some Drift

Reference Mode improves consistency, but it is not a perfect identity lock. For critical brand, legal, or celebrity likeness work, review carefully and use manual finishing where needed.


Example Workflows

Consistent Presenter Clip

  1. Attach 2-4 clear images of the presenter.

  2. Choose a model that supports audio if the presenter should speak.

  3. Prompt the setting, delivery, camera framing, and dialogue.

  4. Review face consistency and lip-sync.

  5. Finish sound and edits in Premiere.

Product Campaign Shot

  1. Attach 2-3 clean product references.

  2. Prompt a new commercial scene or product movement.

  3. Preserve shape, color, label, and material.

  4. Generate multiple versions with different camera or lighting direction.

  5. Upscale or edit the best result if needed.

Character Series

  1. Build a small reference set for the character.

  2. Reuse it across each shot.

  3. Change the prompt for each scene, action, and camera move.

  4. Keep wardrobe and core details consistent unless the story requires a change.


Troubleshooting

Reference Mode does not activate

Make sure Generate Media is enabled, images are attached, no video is attached, and a reference-capable model is selected. If exactly two images trigger Transition Mode, manually switch to a reference model.

The character does not look consistent

Use clearer references, add more useful angles, and include preservation language in the prompt. Avoid mixing references that show different people, outfits, or styles unless that variation is intentional.

The output copies the reference too closely

Use fewer references or add more scene/action detail. The model may be treating your references as the whole shot instead of the identity anchor.

The result ignores the scene prompt

Your references may be too dominant or too visually similar. Reduce the reference count and make the prompt more specific about setting, action, and camera.

The wrong mode activates

Two images often route to Transition Mode. If you want a start/end transition, stay there. If you want identity references, choose Reference Mode manually with a compatible reference model.

The duration or audio options are not what you expected

Reference models have different limits. Some support audio, some do not. Some have fixed duration. Choose the model based on the shot's hardest requirement: audio, duration, reference count, or visual quality.



Next: If your two images are meant to become a start and end frame, use Transition Mode. If you want a guided transition style, use Studio AI Transitions.

Last updated