← AI Video Generation

Image-to-Video AI: A Still-to-Motion Workflow That Keeps Product Details Intact

By Kyle White, Founder of ClipSpeedAIOpen AI Creator →
Summarize this withChatGPTPerplexityGrokGemini
Published October 8, 2026 · Kyle White · 10-minute read

Text-to-video asks a model to invent everything at once: the product, the set, the light, the framing and the motion. When the subject is a real product with a label, a logo or a particular shape, each of those is a chance to get it wrong. Image-to-video splits the job in two. You design one still frame with an image model until it is right, then hand it to a video model whose only job is to make it move.

This guide covers that still-to-motion workflow: building a key frame with an image model such as Nano Banana Pro or GPT Image 2.5, writing a motion prompt, pinning start and end frames, protecting product details, and knowing when reference mode is the better tool. Model facts and prices are as of October 2026 and link to their sources. Check them before you budget.

Pair generated shots with clips from real footage

Image-to-video gives you the product shots you can't film. For the streams, podcasts and long videos you already have, ClipSpeedAI finds the strongest moments, cuts them to vertical 9:16 and burns in captions.

Try ClipSpeedAI →

Why start from a still image

In a standard image-to-video request, your image becomes the clip's first frame. Composition, color, lighting, the product's shape and the label you approved are settled before you spend anything on video, and the video model only decides what changes from there. Stills are also cheap next to video, so fixing a crooked label on the still costs far less than finding it after a render. And one approved key frame can double as the thumbnail, the static ad and the end card.

The workflow has four stages:

  1. Design the key frame with an image model, at the aspect ratio you will publish in.
  2. Make an end frame if the shot has to land somewhere specific, by editing the key frame.
  3. Animate it with an image-to-video model and a prompt that describes motion only.
  4. Check and finish: scrub for drift, add critical text as an overlay, and label it where the platform requires.

Step 1: Design the key frame

Pick the image model for the job

Two current image model families are worth testing for key frames. Our guide to AI image generators for creators compares them in more depth.

Use whichever renders your product more faithfully. If you have real product photos, edit one into a new scene instead of generating the product from a description, so the model works from the real product rather than inventing one.

Compose for the move you plan to make

Decide what the camera and subject will do before you generate the still, then leave room for it:

Treat text and logos as fragile

Even when the still renders a label perfectly, the video model redraws it on every frame after the first, and small type and thin logos have the least margin for error. Make the brand mark large and simple. Keep anything that must be exact, such as prices, claims and legal lines, out of the generated image and add it later as an overlay, where it can't drift.

Step 2: Animate it with a motion prompt

The easiest mistake is re-describing the image. The model can already see the bottle, the slate and the light. Describing them again adds nothing, and if your wording differs from the image, the model has two versions to reconcile. Spend the prompt on what changes: subject motion, one plainly named camera move, the pace, what stays fixed, and sound if the model generates it. Seedance 2.5, Veo 3.1, Kling 3.0 Turbo and MiniMax H3 all generate native audio.

Here is a key frame prompt and a motion prompt for the same shot, with a placeholder where your brand name goes:

Key frame (image model):
Vertical 9:16 product photo. A matte green 500 ml water bottle
with a white label reading "YOUR BRAND" stands on wet slate,
label facing camera, centered in the lower two-thirds.
Soft window light from the left, shallow depth of field,
fine condensation. Plain dark wall, empty space above.

Motion (image-to-video):
Slow push-in over 5 seconds. One drop of water runs down
the side of the bottle. The light warms slightly.
The bottle stays still and the label stays facing camera.
Sound: quiet room tone, one soft drip.

The motion prompt says nothing about color, material, brand or setting, because those are already in the pixels. For prompt structure across models and modes, see the AI video prompt guide.

Start with small moves. A slow push-in reveals little that isn't in the still; a full 360-degree turn makes the model invent the back of the product. If you need that angle, give the model a picture of it, which is what reference mode is for.

Start and end frames: pin both ends of the shot

A start frame fixes where the shot begins and an end frame fixes where it lands; the model generates the motion between them. Use one whenever the last frame matters: a packshot with the logo square to camera, a reveal that ends on the hero angle, or a loop that returns to its first frame. Support varies by model and host:

ModelStart and end framesNotes (October 2026)
Veo 3.1Yes (first and last frame)Clips of 4, 6 or 8 seconds; 1080p, 4K and reference images require 8 (Gemini API docs)
Kling 3.0 TurboYes, on Standard image-to-videoStandard outputs 720p, Pro 1080p (Kling pricing)
Gemini Omni 1.1 FlashYes3 to 10 seconds per clip, extendable to 40 seconds (Google)
MiniMax H3 MaxNot confirmedfal's post-trained H3: 5 to 15 seconds, up to 768p (fal)
Seedance 2.5Not confirmedThe launch post lists image-to-video and reference-to-video; first/last-frame input is unconfirmed and may depend on the host

Make the end frame by editing the start frame, not by generating a second image from scratch. Ask the same image model for "the same image, camera closer, bottle filling two-thirds of the frame" or "the same scene, cap removed and resting beside the bottle". Both image models above can edit, so both frames come from one source image rather than two generations that may disagree on details.

Keep the gap between the frames to something a real camera could cover in the clip's length. If angle, lighting and layout all change at once, the model has to invent a large transition, which is where morphing shows up. One change per shot is a workable rule.

Keeping product details intact

Every frame after the first is generated, so the further a clip travels from your image, in time or camera position, the more the model works from its own idea of the product. Seedance 2.5 can run up to 30 seconds in one pass, which suits narrative work, but for product close-ups, several short shots that each start from their own key frame stay closer to the source.

Product-fidelity checklist
1. Start from a real product photo where you can.
2. Product large in frame, label facing camera and evenly lit.
3. One small camera move per shot; no full rotations without references of the other sides.
4. Say in the motion prompt what stays fixed.
5. Use an end frame to land on the clean packshot.
6. Short shots cut together, not one long take.
7. Prices, claims and legal text go on as overlays, never generated.
8. Scrub the clip frame by frame and compare the last frame with the product photo.

When a detail drifts, change the input before the prompt. A bigger label, a smaller move or a shorter shot reduces what the model has to invent; adding "keep the logo sharp" to the prompt does not.

When to use reference mode instead

In reference mode (reference-to-video in Seedance, Ingredients to Video in Veo), your images guide the model but none of them is the opening frame. The model composes the shot and draws every frame, using your references for what the product, character or setting should look like. Seedance 2.5 accepts up to 30 images, 10 videos and 10 audio clips as references in one request, per its launch post. Veo 3.1's reference images require 8-second clips.

 Image-to-videoReference mode
First frameYour image, exactlyGenerated
CompositionSet by youChosen by the model
New anglesInvented by the modelGuided by extra reference images
Best forPackshots, ad openers, matching a thumbnailProducts in new scenes, multi-angle products, recurring characters

Switch to reference mode when you need the product in scenes you haven't designed, when the shot needs angles your still doesn't show (supply front, back and label close-ups as separate references), or when the same product or character has to appear across several shots. Our guide to consistent AI characters covers the character side.

Stay with image-to-video when the exact opening frame matters or the label must be readable from frame one. The two combine well, too: generate a reference-mode shot to find a composition, take its best frame into the image model, clean it up, and use that as the key frame for an image-to-video pass.

What a still-to-motion iteration costs

The still is the cheap step, so iterate there. These are list prices as of October 2026; the per-clip totals are our arithmetic. For more models, see what AI video generation costs.

StepModel and settingList priceExample
Key frameNano Banana Pro, 1K or 2K$0.134 per image, half in batch10 variations: $1.34
Draft motionGemini Omni 1.1 Flash, 360pAbout $0.03/s (reported)5 seconds: about $0.15
MotionVeo 3.1 Fast, 720p$0.10/s8 seconds: $0.80
MotionKling 3.0 Turbo, 720p$0.112/s5 seconds: $0.56
MotionSeedance 2.5 on fal, 720pAbout $0.473/s5 seconds: about $2.37
FinalVeo 3.1 Standard, 1080p$0.40/s8 seconds: $3.20

The Veo, Kling and Seedance prices include audio. GPT Image 2.5 is left out because OpenAI's own pages list conflicting token rates. Ten candidate stills on Nano Banana Pro cost less than one 8-second Veo Standard clip at 1080p. A cheap draft on a different model is only a rough preview, though: it shows whether the still animates cleanly, not exactly how the final model will move it.

Before you post: labels, rights and provenance

An animated product shot built from a photoreal still can look like real footage, which is what disclosure rules target. YouTube requires creators to disclose realistic synthetic content, including realistic scenes that never happened, using the "AI use" setting in YouTube Studio, and it labels content made with Veo itself (YouTube Help). TikTok requires a label on AI-generated content showing realistic scenes or people (TikTok guidelines). Our guide to AI disclosure rules covers each platform.

Don't assume provenance data follows your still into the video. Nano Banana Pro images carry SynthID, and GPT Image 2.5 adds C2PA metadata and invisible watermarking, but a C2PA chain breaks when a file passes through tools that don't support it, as Google notes. Label the finished video yourself where the platform asks.

Only animate products you are allowed to show, and never generate real people or copyrighted characters without permission. If the clip is an ad, the product on screen should match what customers receive. This is not legal advice; check the current terms of every tool you use.

The bottom line

Image-to-video moves the hard decisions to the cheapest step. Get the still right, ideally starting from a real product photo. Prompt only the motion. Pin an end frame when the shot must land on a packshot, keep moves small and shots short, and switch to reference mode when you need angles or scenes your still can't provide.

ClipSpeed's AI Creator puts these eight models side by side under one account: Nano Banana Pro, GPT Image 2.5 Flare or GPT Image 2.5 Sunburst for the key frame, Gemini Omni 1.1 Flash or MiniMax H3 Max for quick motion tests, then Veo 3.1 for a polished final, Kling 3.0 Turbo for fast variants, or Seedance 2.5 when a shot runs long or leans on references. Generations use ClipSpeed creation credits. It sits next to ClipSpeed's AI clipping, which the ClipSpeed connector also runs from Claude.

Frequently asked questions

What is an image-to-video AI workflow?

You design a single still frame with an image model, such as Nano Banana Pro or GPT Image 2.5, until the composition, lighting and product look right. Then you give that still to an image-to-video model with a prompt that describes only the motion. In a standard image-to-video request your image becomes the first frame, so the hard decisions are made on the cheap step before you pay for video.

Which image model should I use to make the start frame?

Test both families on your own product. Google calls Nano Banana Pro "the best model for creating images with correctly rendered and legible text", which helps with labels and packaging. OpenAI positions GPT Image 2.5 Sunburst for campaign and product imagery, with Flare as the faster default. If you have real product photos, edit one into a new scene instead of generating the product from scratch.

What are start and end frames, and which models support them?

A start frame fixes the first frame of the clip and an end frame fixes the last, with the model generating the motion between them. As of October 2026, Veo 3.1 supports first and last frame, Kling 3.0 Turbo supports it on Standard image-to-video, and Gemini Omni 1.1 Flash takes start and end frames. We have not confirmed it for Seedance 2.5 or MiniMax H3 Max, and support can differ by host.

Why do logos and label text change during image-to-video?

Only the first frame is your image. Every frame after it is generated, so small type and thin logos get redrawn many times as the product or camera moves. Keep the label large and facing camera, use small camera moves and short shots, pin an end frame on the clean packshot, and add prices, claims and legal lines as overlays in your editor instead of generating them.

When should I use reference mode instead of image-to-video?

Use reference mode when you need the product in scenes you haven't designed, when the shot needs angles your still doesn't show, or when the same product or character must appear across several shots. Seedance 2.5 accepts up to 30 reference images per request. Stay with image-to-video when the exact opening frame matters or the label must be readable from frame one.

How much does it cost to animate a still image with AI?

At October 2026 list prices, a Nano Banana Pro still is $0.134 at 1K or 2K. By our arithmetic, an 8-second Veo 3.1 Fast clip at 720p is $0.80, a 5-second Kling 3.0 Turbo clip at 720p is $0.56, and a 5-second Seedance 2.5 clip at 720p on fal is about $2.37. Iterate on the still, where mistakes are cheapest.

Can I do image-to-video inside ClipSpeed?

Yes. ClipSpeed's AI Creator puts both halves of this workflow in one workspace: Nano Banana Pro, GPT Image 2.5 Flare and GPT Image 2.5 Sunburst for the key frame, and Veo 3.1, Kling 3.0 Turbo, Gemini Omni 1.1 Flash, MiniMax H3 Max and Seedance 2.5 to animate it. Generations use ClipSpeed creation credits; see pricing. The same account also does AI clipping, which turns long videos and live streams into captioned vertical clips.

Related guides

Published by ClipSpeedAI · AI video generation and AI clipping in one place — create with Seedance, Veo, Kling and Nano Banana, then cut it into captioned shorts.