For the complete documentation index, see llms.txt. This page is also available as Markdown.

Image-to-Video

Animate still images into motion by uploading an image and describing the movement you want. Perfect for bringing static artwork, photos, or designs to life.

Image-to-Video turns a still image into a moving shot. You provide the first frame, then describe how the camera, subject, and environment should move.

Use it when the image already has the visual direction you want: a product photo, generated frame, concept image, character portrait, landscape, frame grab, title-card design, or artwork. The model does not need to invent the entire scene from text. It starts from your image and tries to bring it to life.

If your main goal is a controlled camera movement like dolly in, orbit, crane up, tracking shot, or handheld, use Studio Motion Director. It is a guided Image-to-Video workflow built specifically for camera moves.


When To Use Image-to-Video

Use Image-to-Video when you already have the frame and need motion:

  • Animate a product image.

  • Bring a Cinematic Lab frame to life.

  • Turn concept art into a short moving insert.

  • Add movement to a landscape, environment, or real estate still.

  • Animate a character portrait or illustration.

  • Create motion from a Premiere frame capture.

  • Test whether a still can become usable B-roll.

Choose another workflow when:

You have...
Use instead

Only a text idea

Text-to-Video

A still image and want guided camera presets

Studio Motion Director

Two images to connect

Transition Mode or Studio AI Transitions

Several references for the same character/product

Reference Mode

A still image that needs to be created first

Studio Cinematic Lab

Existing footage to change

Studio


How It Works

  1. Enable Generate Media in the composer.

  2. Attach one image.

  3. Choose a model that supports Image-to-Video.

  4. Set aspect ratio, duration, resolution, and audio options.

  5. Write a motion-focused prompt.

  6. Generate the video.

  7. Review the motion and regenerate, edit, or send the result into Studio.

The image provides the visual anchor. Your prompt should describe motion, not re-describe the picture.


Start With A Strong Image

The source image matters more than the model. A weak still usually becomes a weak video.

Good Image-to-Video sources have:

  • A clear subject.

  • Readable composition.

  • Enough room for motion.

  • A stable aspect ratio that matches the final video.

  • Visible depth: foreground, subject, and background.

  • Good lighting and texture.

  • No important details cut off at the edge.

Risky source images include:

  • Tiny subjects in a busy frame.

  • Faces or hands already distorted.

  • Important text or logos that must remain perfect.

  • Flat graphics with no depth.

  • Cropped portraits with no room for camera movement.

  • Images with multiple subjects when only one should move.

If you need a better still before animating, create one in Studio Cinematic Lab first.


Write Motion Prompts, Not Image Prompts

The image already tells the model what the scene looks like. Your prompt should tell it what changes over time.

Focus on:

Motion layer
What to describe

Camera movement

Push in, pull out, pan, orbit, crane, handheld, tracking.

Subject movement

Turns, walks, looks up, lifts product, hair moves, fabric shifts.

Environment motion

Wind, water, clouds, light flicker, particles, crowd movement.

Mood and pacing

Calm, energetic, elegant, tense, dreamy, documentary.

Preservation

What should stay stable or unchanged.

Useful structure:


Good Prompt Examples

Product Animation

Portrait Motion

Landscape Motion

Artwork Or Poster Motion

Character Illustration


Weak Prompts To Avoid

Too static:

Better:

Too vague:

Better:

Too much change:

Better:

Image-to-Video works best when the prompt respects the still image instead of trying to replace it.


Aspect Ratio And Framing

Match the video aspect ratio to the source image whenever possible.

Source image
Best video aspect ratio

Landscape frame

16:9

Vertical portrait or social frame

9:16

Square design or product post

1:1

Cinematic wide frame

21:9, if supported by the selected model

Why this matters:

  • It prevents unexpected cropping.

  • It keeps the first frame close to your original image.

  • It preserves product, face, and composition placement.

  • It gives the model a stable visual anchor.

If the source image is too tight, generate or crop a wider version first. Image-to-Video needs room to move.


Choosing A Model

Use the model based on the hardest part of the animation.

Need
Good starting point

Dialogue or audio from a still

Veo 3.1, Veo 3.1 Fast, or Google Omni Flash

Budget visual drafts without audio

Veo 3.1 Lite or Seedance 2 Mini

Strong camera movement

Kling 3.0 Pro, Kling O3, or Studio Motion Director

Longer cinematic visual clips

Sora 2 or Sora 2 Pro

Natural movement with ambient audio

Seedance 2

High-energy action

Hailuo 03

2K from a still, with audio

Hailuo 03

The same animation back in seconds

Hailuo 03 Max (480P/768P, near-realtime, audio on)

Fast animation with lipsync

Grok Imagine 1.5

Flexible output with audio, 2–30 seconds

Wan 3.0 (480p/720p/1080p, audio on by default)

A clip of 15–20 seconds at 1080p

Flux 3 (up to 20s)

A clip that runs past 20 seconds

Seedance 2.5 (up to 30s, up to 1080p)

Ultra-wide 2:1 or 21:9 framing

Flux 3

A cheap, silent 5s animation or start/end pair

Luma Ray 3.2 (540p–1080p)

For a deeper chooser guide, see Supported Video Models.

Flux 3 animates a still for up to 20 seconds with native audio at 720p or 1080p (1080p by default). Its aspect ratio defaults to Auto, following your source image, but unlike Hailuo 03 it is not locked to the source: you can override with any of eight ratios, including 2:1 and 21:9.

Two things to know before you pick it. First, once you attach an image, Flux 3 offers only its image-to-video tier — the text-to-video tier hides, which is stricter than Kling. (The Hailuo tiers hide their text-to-video entries at one image too, but Hailuo keeps its reference tiers where Flux 3 has none.) Remove the image to get text-to-video back. Second, Flux 3 has no reference tier, so at three or more images it disappears from the picker entirely and a reference model such as Veo 3.1 Reference takes over.

Cost scales with duration: $0.17/second at 720p, $0.29/second at 1080p. A 20-second 1080p animation is $5.80.

Seedance 2.5 animates a still for 4–30 seconds, or lets the model choose. The duration pill runs 4 to 30 whole seconds plus Auto, which is the default and lets Seedance pick the length from your prompt. Output is 480p, 720p, or 1080p (default 720p, with no 4K) and carries native audio you can switch off. Attach a second image and it becomes the end frame on the same endpoint — nothing re-routes, the way Luma and Hailuo work rather than the way Flux 3 does.

Unlike Flux 3 and Luma, attaching an image does not narrow the Seedance group to one tier: at one or two images you still see the main tiers and the Reference tiers together. At three or more images the group narrows to the Reference tiers, and a Seedance 2.5 escalation lands on Seedance 2.5 Reference, never on a Seedance 2 tier.

Billing is token-based, so cost per second depends on the aspect ratio: roughly $0.4730/second at 720p 16:9, $0.2205 at 480p, and ≈$1.04 at 1080p 16:9 (1080p rate ⚠️ UNVERIFIED — confirm in the Fal playground before release), which puts a 30-second 720p animation near $14.

Hailuo 03 Max animates a still in near-realtime. It runs 5–15 seconds at 480P or 768P (768P by default) with audio always on, and comes back in roughly the time the clip itself runs — where base Hailuo 03 renders 2K and can queue for a long while. Attach a second image and it becomes the end frame on the same endpoint, exactly as on base Hailuo 03; nothing re-routes and there is no separate transition entry.

Like the base tier, Max image-to-video has no aspect-ratio control — the output follows the image you attach, so crop the still first if you need a different shape. Attaching an image hides both Hailuo text-to-video tiers, but the family stays in the picker: at one or two images you see both image-to-video tiers alongside Hailuo 03 Reference and Hailuo 03 Max Reference.

At about $0.08 per second at 768P ($0.05 at 480P), a 10-second animation is roughly $0.80 — cheap enough to try several motion prompts before re-rendering the keeper on Hailuo 03 at 2K.

Luma Ray 3.2 animates a still for exactly 5 seconds, silently. The image-to-video tier is 5s only (10s needs multi-keyframe input that Chat Video Pro does not expose), renders at 540p, 720p, or 1080p (default 720p), and generates no audio on any tier. Attach a second image and it becomes the end frame on the same endpoint — unlike Flux 3, nothing re-routes to a separate first/last-frame endpoint. Luma narrows the same strict way Flux 3 does: with an image attached, only the image-to-video tier is offered, and at three or more images the family leaves the picker (no reference tier). At the 720p default a 5s animation costs $0.30 ($0.06/second; $0.03 at 540p, $0.24 at 1080p).

Google Omni Flash animates a still for 3–10 seconds, with audio always on. Output runs at 360p, 720p, 1080p, or 4K, and Chat Video Pro sends 1080p by default. There is no audio pill in the family — sound is generated with the picture on every clip, so describe it in the prompt.

Attach a second image and it becomes the end frame on the same endpoint, interpolating the first frame into the second — the way Luma and Seedance 2.5 work rather than the way Flux 3 does. This changed with the v1.1 upgrade: two images used to promote the generation to Reference mode automatically. If you want the pair treated as references instead, choose Omni Flash Reference from the submenu, which now accepts up to 10 images plus up to 3 reference clips of 3 seconds or less.

Cost follows the resolution pill: $0.03/second at 360p, $0.10 at 720p, $0.15 at 1080p, and $0.30 at 4K — so a default 8-second 1080p animation is about $1.20.


Image-to-Video vs. Studio Motion Director

Both workflows animate still images, but they are not the same.

Use Image-to-Video when...
Use Studio Motion Director when...

You want direct model control.

You want guided camera movement presets.

You already know the exact model to use.

You want the workflow to build the movement prompt.

The motion is custom or unusual.

The motion is a known shot type: dolly, orbit, crane, tracking, handheld.

You want to test several models manually.

You want a more structured, production-style camera move.

Motion Director is usually better for users who think in shot language. Image-to-Video is better when you want direct control in the model selector.


Best Practices

Keep The First Move Simple

Start with one clear motion. Add complexity only after the first result works.

Good:

Riskier:

Preserve Important Details

If something must stay stable, say so:

This is especially important for products, faces, logos, packaging, and character art.

Use Depth

Image-to-Video benefits from images with foreground, subject, and background separation. Depth gives the model something to animate through parallax.

Good sources:

  • Product on a table with background props.

  • Portrait with lights behind the subject.

  • Landscape with foreground plants and distant mountains.

  • Interior scene with layers: doorway, subject, background window.

Do Not Ask For A New Scene

If the prompt asks for a totally different location, outfit, or subject, the model may fight the image. Use Text-to-Video or generate a new still in Cinematic Lab instead.

Review The First And Last Frames

A result can look good in the middle but drift at the beginning or end. Check whether the first frame still matches your image and whether the final frame stayed coherent.


Example Workflows

Animate A Cinematic Lab Frame

  1. Create a strong still in Cinematic Lab.

  2. Use Image-to-Video or Motion Director.

  3. Prompt for one clear camera movement.

  4. Review identity, framing, and motion.

  5. Use Upscale after the creative result is approved.

Product Motion Shot

  1. Start with a clean product image.

  2. Match the output aspect ratio to the source.

  3. Prompt for a slow orbit, push-in, or studio-light movement.

  4. Preserve product shape, label, and color.

  5. Use the result as B-roll or ad footage.

Social Image Animation

  1. Start with a vertical 9:16 image.

  2. Use a model that supports the needed vertical format.

  3. Prompt for bold, readable motion.

  4. Keep the action simple for mobile viewing.

  5. Add captions, music, or voiceover in Premiere.


Troubleshooting

The image barely moves

Make the motion more specific. "Slow push-in" or "camera pans left to reveal the background" is stronger than "animate this image."

The motion is too extreme

Use gentler language:

The subject changes identity

Add preservation language and simplify the action:

The image is cropped

Match the output aspect ratio to the source image. If you need a different format, create a version of the still in that format first.

The model ignores part of the prompt

Reduce the number of instructions. Image-to-Video is easier to control when the prompt has one main camera move, one subject action, and one environmental motion.

The result is close but needs polish

Use Studio for the next pass:

Problem
Better next step

Needs a more controlled camera move

Motion Director

Needs a different angle first

Multi-Cam

Needs better lighting

Relight Scene

Needs VFX or atmosphere

Add Effects

Needs resolution

Upscale



Next: If you need a directed camera move from your still, use Studio Motion Director. If you need to connect two frames, use Transition Mode or Studio AI Transitions.

Last updated