Image-to-Video: Animate a Still You Already Like
The cheapest reliable route to good AI video on Gendia, generate a still for 20 credits, then make it move. How first and last frame inputs work, and what animates well.
Text-to-video asks a model to invent a composition, a subject, a look and a motion, all at once, at video prices. Half of what you're paying for is a picture you can't see until it's finished.
Image-to-video splits that in two. Get the picture right for 20 credits where mistakes are cheap. Then pay for the motion only.
It's the single biggest improvement most people can make to both their results and their bill.
Why it's more controllable
When you attach a still, the composition, subject, colour, lighting and framing are all already decided. The model isn't guessing at any of it. Its entire job is movement.
That means:
- You know exactly what the first frame looks like, because you chose it
- Your product looks like your product
- The look is consistent across a whole set of clips
- The prompt has one job, so it can be short and specific
Here is the whole idea in two steps. First the still: one Seedream generation, 20 credits, iterated until the light and the composition were right:
Then the same picture handed to a video model as its first frame, with a prompt that talks only about movement:
The bottle in the clip is the bottle from the still, because it was never re-invented. That is the entire argument for working this way.
First frame, last frame
Some models take two images: the start and the end, and generate the motion between them.
This is the most controlled form of AI video available. You decide both ends; the model only fills the middle.
It's ideal for:
- A product rotating or opening, supply both positions
- A transition between two states, day to night, before to after, closed to open
- Matching an ending frame to the next clip's beginning, so a sequence cuts cleanly
Where the two frames are close in composition, the result is usually excellent. Where they're wildly different, the model has to invent a transformation, and you get a morph rather than a move.
What animates convincingly
- Slow camera moves on a static subject. A push in on a product. Almost always works.
- Natural ambient motion. Steam, liquid, fabric in a breeze, hair, foliage, water, firelight. Models handle these beautifully and they add life without risk.
- A single simple subject action. Someone turns their head, a hand lifts a cup.
- Light changing. A sweep across a surface, a shadow moving.
What tends to break
- Fast or complex human movement. Walking is manageable; dancing, sport and gesture-heavy action deform.
- Hands doing anything precise.
- Crowds. Many small figures moving means many small failures.
- Legible text in the frame. Text is stable in a still and rarely stable in motion: it warps as the frame changes.
- Anything requiring the model to invent what's behind the subject, big camera moves reveal areas that weren't in your still, and it makes them up.
The full cheap workflow
This is how I actually make a product clip, end to end:
1 · Generate the still on Seedream 4.5, 20 credits. Iterate here, cheaply, until the composition and light are right. Three or four attempts is 80 credits.
2 · Test the motion on LTX Video, 5 seconds, 480p, 20 to 60 credits. Does the movement idea read? Does the product hold together?
3 · Render properly on the model the job deserves, once.
Roughly 150 credits spent before the expensive render, and you arrive at it knowing precisely what you're buying. The alternative, writing a text-to-video prompt straight onto an expensive model, costs more per attempt and gives you less to learn from, because when it's wrong you don't know whether the picture or the motion was the problem.
Prompting when a still is attached
Don't describe the picture. It's already there. Describing it again just competes with it.
Bad: "A white ceramic bottle on a marble surface in soft light", with that exact image attached.
Good: "Slow push in. Steam rises gently from the left. Light sweeps slowly across the label."
Motion, camera, pace. Nothing else.
Try it
Take any image you've already generated, open the video generator, attach it on LTX at 5 seconds and 480p, and write one line of motion. Under 60 credits to watch your own picture move.
- #ImageToVideoAI
- #AnimateAPhotoAI
- #AIPhotoToVideo
- #FirstFrameLastFrameVideo
- #MakeAPictureMoveAI
- #AIProductVideo
- #StillImageAnimationAI
- #AIVideoFromImage
- #GendiaTutorial
- #Gendia



