Name the shot, not the feeling: film's shot vocabulary tells a model exactly where the camera sits, and it still works when a model release breaks everything else.
- Angle is the most reliable lever. Low angle reads as power, high angle as vulnerability.
- One of each, no more. Pick one angle, one distance and one movement; stacking terms averages the shots.
- Movement needs time. A slow push-in is meaningless in a two-second generation.
- Near-synonyms differ. Handheld feels urgent and human; shaky means loss of control.
- When stuck, simplify. Composition and angle are understood more reliably than movement.
Twenty-four shot types, grouped by what they actually control. This vocabulary predates AI by a century — which is exactly why it still works when a model release breaks everything else.
A generative model has seen millions of films. It knows what "low angle" means far better than it knows what "make it look powerful" means.
Naming the shot is not jargon. It is the difference between describing a feeling and specifying a camera position.
Angle — where the camera sits
The single most reliable lever, because it maps to how humans already read physical height and dominance.
Overhead and high angle are not the same note. High angle diminishes a person; overhead removes them from the situation entirely — they become one element in a pattern.
Where the camera sits sets power. Low angle makes a subject look powerful, high angle vulnerable, and overhead removes them from the situation entirely.
Distance — how much is in frame
Close-up carries a range rather than one meaning — intimacy when the subject is open, intensity when they are not. The distance sets the access; the performance decides what you do with it.
Distance sets how close the audience gets. A close-up can mean intimacy or intensity; the performance decides which.
Movement — what the camera does during the take
Push-in and long take both produce tension by different means. One approaches; the other refuses to release. Knowing which you want is the difference between a shot that builds and one that merely lasts.
Handheld and shaky are also not synonyms. Handheld implies a person holding the camera — urgent, present, human. Shaky implies loss of control. Prompting for one and getting the other is a common and fixable failure.
On anamorphic: it names a physical lens, and its signatures — oval bokeh, horizontal blue flares, a wider frame — are specific enough to ask for individually. The realism words sets out what the lens actually does.
Pick movement by the effect you want. A push-in builds tension by approaching, a long take by refusing to cut, and handheld is not the same as shaky.
Focus — what is sharp, and when
These three control when the audience learns something, which is why they are the hardest to specify and the most valuable when they land.
Focus controls when the audience learns something. Rack focus, POV and foreground reveals are the hardest to specify and the most valuable when they land.
Composition — where the subject sits
Centre frame and locked shot both read as "control" — and they are not the same thing.
Locked is camera behaviour: nothing moves. Centre frame is composition: the subject owns the middle. Combine them and the control reads as absolute. Use one against the other — a locked camera on an off-centre subject — and you get something considerably more unsettled.
Where the subject sits in frame carries meaning. Centre frame and a locked camera both read as control, but a locked camera on an off-centre subject feels unsettled.
Using this in a prompt
- Name the shot, not the feeling. "Low angle" outperforms "make him look powerful" because it specifies a camera position rather than an outcome.
- One angle, one distance, one movement. Stacking five terms produces an average of five shots.
- Movement needs duration. "Slow push-in" is meaningless in a two-second generation.
- Composition survives model changes better than movement does. If a generation keeps failing, drop to composition and angle — those are the most reliably understood.
The full prompt structure is on anatomy of an AI video prompt, and current model capability is on the models page.
Name the shot, not the feeling. Use one angle, one distance and one movement, and fall back to composition and angle when a generation keeps failing.