Prompting for Video: Motion, Camera and Light That Actually Land
A prompt that produces a beautiful still frequently produces a disappointing clip. The reason is structural rather than mysterious: an image prompt describes a state, and a video prompt has to describe a state and a change. Leave the change out and the model invents one, usually a drifting camera and a subject that melts.
This guide covers what to write instead. It assumes you have run a first clip and have a shot list to work from.
The rule that fixes most failures
Start from a still, not a description.
There are two ways into a video generation. Writing a description and generating from nothing gives the model authority over composition, subject, light and motion all at once. Supplying an approved still and describing only the motion gives it authority over one thing.
The second is better most of the time, and it is not close. It is also cheaper, because settling composition on stills costs a fraction of settling it on clips. The working habit is: generate a still you like, approve it, then animate that.
Use a description-only video generation when you genuinely do not care what the shot looks like — an establishing texture, a background plate — and want the fastest route to something usable.
The four things a video prompt has to carry
| Element | Question it answers | Example |
|---|---|---|
| Subject motion | What do the things in the shot do? | she turns her head slowly toward the camera |
| Camera motion | What does the camera do? | camera holds steady |
| Continuity | What must not change? | red wool coat, night, neon signs |
| Pace | How fast? | slowly, gently, sharply |
Drop any one and the model fills it in. Camera is the one people drop, and it is the one that most visibly ruins a shot.
Say what the camera does, every time
If you do not mention the camera, you get whatever the model decides — and models decide in favour of movement. Unmotivated drift is the single most common reason a generated clip looks amateur next to a real one.
“Camera holds steady” is the most useful phrase in video prompting. Real drama is full of locked-off shots; generators are not, unless told.
The vocabulary worth knowing:
- camera holds steady — no movement. Correct far more often than instinct suggests.
- slow push in / slow pull out — moving toward or away. Push in builds tension; pull out reveals context.
- pan left / pan right — pivoting in place.
- tracking shot, follows her — moving with the subject.
- handheld, slight shake — deliberate instability. Use sparingly; it is a strong flavour.
One camera instruction per shot. Combining a push in with a pan and a handheld shake asks for three things and usually gets a mess.
Ask for one action
A clip a few seconds long cannot deliver a sequence of events. This is the hardest habit to build, because writing three things feels more productive than writing one.
| Asking for too much | What to write instead |
|---|---|
| She walks in, sees the letter, and drops her keys | She stops in the doorway, eyes fixed on the table |
| The car pulls up, the door opens, he steps out | The car door swings open, dust settling around it |
| Rain starts, lights flicker, she turns | Rain streaks the window as she turns her head |
The rejected column is not lost — it is three shots instead of one, which is what the shot list is for.
Motion has a speed, and the default is too fast
Generators tend toward brisk. Naming the pace is a one-word fix that makes clips read as directed:
she turns her head slowly toward the camera, rain continues to fall,
camera holds steady
Slowly and continues are doing real work there. Continues is especially useful for ambient motion — rain, smoke, steam, fabric — because it tells the model the movement is ongoing rather than an event.
Describe the result, not the picture
When you animate from a still, the model can already see the picture. It cannot see your intention.
Given a frame of a kitchen at night, do not write “a kitchen at night”. Write what should be different a second later.
The same principle governs the strength control when you are reworking a still with img2img — the setting that decides how much changes:
| Setting | What you get |
|---|---|
| Low | Almost your original, lightly touched |
| Medium | Clearly the same scene, genuinely reworked |
| High | Your picture as a loose suggestion |
Two failure modes, two fixes: nothing seems to have changed means strength is too low; your subject has vanished means it is too high. Start in the middle and move in small steps.
Say what you want, not what you do not
“A street with no cars” often produces a street with cars, because the description still contains the word cars. Describe the thing you do want — “an empty street at dawn”.
If the panel offers a negative prompt box, that is the correct place for things to avoid. Putting them in the main prompt tends to summon them.
Length and ordering
Two or three lines is the sweet spot: enough for subject, motion, camera and continuity; short enough that nothing gets lost.
Very long prompts are not more powerful. Past a point, later words get less attention than earlier ones, so a hundred-word prompt often ignores its own ending. Put what matters most first — usually the action.
Reference material
The video tools accept reference material as well as a prompt: up to nine images and three videos, plus up to three audio clips. Resolution options run from 480p to 4K.
Two distinctions worth keeping straight:
- A source image is the thing being changed. Your approved still.
- A reference image is guidance for making something new — keeping a subject or a style consistent, rather than transforming that picture.
The field says which it is. Confusing them is a quiet source of results that ignore either your picture or your prompt.
Larger resolutions cost more and take longer. Generate at a normal size and upscale the shots you keep.
Diagnosing a bad clip
| Symptom | Cause | Fix |
|---|---|---|
| Camera wanders around | Camera unspecified | Add “camera holds steady” |
| Subject warps or melts partway | Motion too complex | Ask for simpler movement |
| Nothing much happens | Description, not direction | Name the action explicitly |
| Result ignores my still | Strength too high | Lower it |
| Result ignores my prompt | Strength too low | Raise it, and check the prompt describes the change |
| Faces drift across shots | No locked character | Use an avatar |
| Clip shorter than expected | Model-dependent length | Check what that model supports |
| Nothing looks like my idea | Started from a description | Start from a still instead |
More failure modes, including queue states and partial batch failures, are covered in when a generation fails.
Change one thing at a time
The habit that compounds. Generate, change one element, generate again.
Change the motion, the camera and the light at once and get a worse result, and you have learned nothing about which change did it. Change one and you learn something you can apply on every shot afterwards — which is what turns a first week of guessing into a second week of production.
Keep the prompts that work. A prompt that reliably produces the look you want is worth more than any individual clip, because it applies to a hundred different shots.
What to read next
- What an episode actually costs — drafting cheap is a prompting decision as much as a budget one.
- From clips to a finished episode — where these shots go.
- Producing a season, not a clip — running a proven prompt at volume.
Last reviewed: · Editorial policy · Report an error