Prompt Writing for Motion: Why Your Clips Look Like Wobbling Photographs
Video prompts need a different structure from image prompts. How to describe camera moves, subject action and pacing on Gendia so a clip does something instead of just existing.
Your first video prompt was probably a good image prompt. Mine was. And what came back was a beautiful still photograph with a faint, aimless drift: the AI video equivalent of a screensaver.
The fix isn't more description. It's describing a different kind of thing.
An image prompt describes a moment; a video prompt describes a change
That's the whole idea, and everything below follows from it.
Image prompt: "A woman in a red coat standing on a rainy street at night, neon reflections."
Video prompt: "A woman in a red coat walks toward the camera down a rainy street at night, neon reflections shifting across the wet pavement as she passes. Slow push in. Steady pace."
Same scene. The second one tells the model what should be different at the end than at the start, which is what a video is.
If your prompt would work unchanged as an image prompt, it will give you a wobbling photograph.
Here is that difference, generated on Seedance 2.5, same scene, same model, same five seconds, only the prompt changed.
The first clip is not broken. It is what a model does when you give it a photograph to make and five seconds to fill: something, gently, in no particular direction. The second one was told what to do.
The four things worth saying
1 · The subject and setting. Same as an image prompt, and keep it short: this is not where the value is.
2 · The action. What actually happens. Someone walks, liquid pours, light sweeps, a door opens. One action, described plainly.
3 · The camera. What the camera does, separately from what the subject does. Static? Pushing in? Orbiting?
4 · The pace. Slow, steady, sudden. Left unsaid, most models default to a gentle mid-speed drift that reads as nothing in particular.
Camera vocabulary that works
These terms come from film and the models understand them:
- Static shot, locked off, no camera movement. Genuinely underrated: it makes the subject's action the whole story, and it's the most reliable thing to ask for.
- Slow push in / dolly in, moving toward the subject. Builds attention. The workhorse of product video.
- Pull back / dolly out, reveals context.
- Pan left / right, pivoting in place.
- Tilt up / down, pivoting vertically. Good for scale.
- Tracking shot, moving alongside a subject.
- Orbit, circling. Impressive on products, demanding on the model.
- Handheld, deliberate slight instability, for documentary feel.
- Crane / drone shot, large sweeping moves.
Say only one. "Slow push in while panning left and orbiting" gives you incoherent motion. Pick the move that serves the shot.
And know when to ask for nothing: if the subject's own action is interesting, a static camera is usually stronger than a moving one.
One action per clip
The single most common reason a clip looks like mush.
Too much: "She walks into the room, sets down her bag, takes off her coat, sits, opens a laptop and begins to type."
That's five actions. In five seconds you'd get half of each, blurred together.
Right: "She sets her bag down on the table and looks up."
If you need the whole sequence, that's several clips assembled in the timeline editor, which is also how real film works.
Where clips actually go wrong
Nothing happens. No action described, so the model adds ambient drift. Fix: name an action.
Everything happens. Too many actions in too few seconds. Fix: cut to one.
The subject morphs. Faces and objects deform mid-clip, usually when the motion asked for is too fast or too complex. Fix: slow the pace, simplify the move, or supply a reference image.
The motion is right but the shot is dull. The action works and the camera is doing nothing interesting. Fix: add one camera move, usually a slow push in.
It's fine but too fast. Models tend to over-animate. "Slow", "gentle", "gradual" are the most useful words in video prompting.
The cheap testing loop
Never debug a prompt on an expensive model. The routine:
- LTX Video, 5 seconds, 480p, 20 to 60 credits.
- Rewrite until the motion idea works. Five attempts is around 200 credits.
- Move the finished prompt to the real model at the real duration and resolution, and run it once.
Compare that with debugging on Seedance 2.5 at 1080p, where five attempts at ten seconds costs over 16,000 credits. Same five attempts, same learning, wildly different bill.
Sound, on models that generate it
Where audio is generated with the picture, Seedance 2.5 among them, describe what you want to hear alongside what you want to see. "Rain on the pavement, distant traffic, no music" is a valid part of a video prompt and it will be respected.
Try it
Open the video generator, pick LTX, and write the same scene twice: once as an image prompt, once with an action, a camera move and a pace. Under 120 credits, and the pair will teach you more than any list of tips.
- #AIVideoPromptGuide
- #VideoPromptEngineering
- #CameraMovementPrompt
- #AIVideoTips
- #HowToPromptAIVideo
- #TextToVideoPrompt
- #CinematicAIVideo
- #AIVideoMotion
- #GendiaTutorial
- #Gendia



