ONLINEAGENT_OPS 2026.Q3 HOME ARTICLES CRAFT RECORD BLOG MAP HUBS FAQ SEARCH
HOMEARTICLESThe Anatomy of an AI Video Prompt
ARTICLES · CRAFT

The Anatomy of an AI Video Prompt

How to structure an AI video prompt that works. Bracket format, motion-first ordering, identity lock, camera specification — built from real Seedance.

READ4 min
WORDS1,022
SECTIONS7
TYPEEXPLAINER
CHECKED25 AUG 26
TL;DR — THE SHORT VERSION

A good video prompt names the motion, locks identity, sets the camera and states duration; use brackets as a checklist, then order it the way your vendor asks.

  • Vendors differ on order. Seedance 2.5, Runway, Veo 3.1 and Kling 3.0 each publish different prompt structures.
  • Brackets are a checklist. They make sure motion, identity, camera and duration are all covered.
  • Lock identity in words. Add "no morphing" to the subject, plus a reference photo for high-movement shots.
  • Keep the camera separate. Mixed instructions can make the subject move instead of the camera.
  • Always state duration. Add "no scene cuts", or some models insert a hard cut mid-clip.

AI video prompting is more demanding than image prompting. You're not describing a single frame — you're describing motion, time, and change. Every element that can go wrong in an image can also drift from frame to frame across a clip. Structure is the most reliable tool you have for keeping it consistent.

This article covers the structure that works: the bracket format, the motion-first principle, identity lock, camera specification, and duration.

THE BRACKET FORMAT IS NOT MODEL-AGNOSTIC — CORRECTED 25 AUG 2026

An earlier version of this page said the same structure applies to every video model. Checked against the vendors' own documentation, that is not true. They publish different structures — set out side by side on what the vendors actually say — and none of them is this bracket format.

Seedance 2.5ByteDance publishes an eight-part formula, and it is not this one. Through Dreamina, their own platform, the stated order is Format → Subject → Action → Environment → Camera → Look → Timing → Audio & Constraints. Format comes first (duration, aspect ratio, shot structure) and action comes third, not first. Their template also binds references with @Image1/@Image2, puts shot size before style language, allows one camera move per shot, budgets roughly 6–8 seconds per beat, and places exclusions last — after a positive Preserve [identity, wardrobe, logo, object geometry] clause. See the full template.
Runway Gen-4/4.5no fixed order. Runway: “Prompts do not need to follow a specific structure in most cases to receive quality outputs, but we recommend prompting with full sentences for more control over certain elements and appending additional keywords in cases where more variation is welcomed.” For video it says to start from “only the most essential motion to the scene” and add elements one at a time. An earlier version of this line gave Runway an order ending in motion over time and style, and said a bracket prompt works against it; neither is in Runway’s guides.
Veo 3.1 — Google's published formula is [Cinematography] + [Subject] + [Action] + [Context] + [Style & Ambiance], which opens with the camera, not the motion — the reverse of the motion-first principle below. Veo also has its own audio syntax: dialogue in quotation marks, effects prefixed SFX:, background prefixed Ambient noise:.
Kling 3.0 — Kling’s own formula is Subject + Subject Movement + Scene. Its separate negative-prompt field exists only on the legacy API, and Kling recommends writing negatives inside the prompt anyway. An earlier version of this line called it a real negative-prompt field.

What carries across better than the syntax is the thinking: name the motion explicitly, anchor identity so it cannot drift, specify the camera once, and state duration. The vendors disagree on order and packaging.

So what is the bracket format below? It is a working structure that was tested on Seedance 2.0 in June 2026 and produces good output. It is not, and this page should not have implied it was, the structure ByteDance publishes. Use it as a checklist that guarantees you have not forgotten motion, identity, camera or duration — then order the finished prompt the way your model’s vendor asks. Per-vendor templates with sources: what the vendors actually say.

Vendor documentation: Luma and Melies Seedance guides · Runway, Gen-4 Image and Gen-4 Video prompting guides, re-read in a browser 17 Sep 2026 · Google Cloud Veo 3.1 prompting guide · Kling API docs. Checked 25 Aug 2026

Why Video Prompts Need More Structure Than Image Prompts

In a single image, a vague subject description produces a generic image. In a video, a vague subject description produces a different-looking person in every frame. The model isn't just guessing the subject once — it's regenerating its interpretation of the subject across every frame, and without explicit anchoring it will drift.

Motion introduces another layer. The model needs to understand what's moving, how it's moving, at what speed, and from what camera angle — all simultaneously. Without clear structure, motion instructions bleed into subject descriptions, camera movements get conflated with subject movements, and the output becomes incoherent.

TAKEAWAY

Video drifts frame by frame. A vague subject can look different in every frame, and loose motion instructions bleed into subject and camera descriptions.

The Bracket Format

The bracket format organises video prompt elements into labelled sections. Labelling each element keeps motion, subject and camera instructions apart, and makes it easy to see that nothing is missing. No vendor documents that a model reads the labels as separate sections.

[MOTION]: describe what moves, how it moves, and at what pace [SUBJECT]: who or what is in the scene + identity lock instructions [ENV]: the environment, lighting, and atmosphere [CAMERA]: camera type, movement, lens, and shooting style [STYLE]: light, color grade, film aesthetic [DURATION]: length in seconds, frame rate, and cut instructions

You don't need all six sections for every prompt. Short lifestyle clips might only need [MOTION], [SUBJECT], [ENV], and [DURATION]. Complex cinematic sequences need all of them. The key is: whatever you include, put it in its own labelled bracket.

SAME INGREDIENTS · FOUR ORDERS
This page’s bracket format beside the three vendor structures in its correction box, each read top to bottom. Motion or action is highlighted: first here and in Runway’s, third in Seedance’s and Veo’s.
BRACKET FORMAT[MOTION][SUBJECT][ENV][CAMERA][STYLE][DURATION]SEEDANCE 2.5FormatSubjectActionEnvironmentCameraLookTimingAudio & constraintsRUNWAY GEN-4/4.5Essential motionSubject motionCamera motionScene motionStyleVEO 3.1CinematographySubjectActionContextStyle & ambiance
Reasoning — summarises this page’s correction box and bracket format; the vendor orders are the ones that box gives from vendor documentation checked 25 Aug 2026; Runway’s guides re-read in a browser 17 Sep 2026.
TAKEAWAY

Give each element its own labelled bracket. You do not need all six sections; short lifestyle clips may only need motion, subject, environment and duration.

The Motion-First Principle

This page’s bracket format puts motion first, so the one thing a still image cannot show is never left out. Runway’s video guide starts there too: “Begin with a foundational prompt that captures only the most essential motion to the scene.” Other vendors order it differently, and no vendor documents that video models in general weight earlier words more heavily, so treat motion-first as this format’s habit rather than a rule.Runway, Gen-4 Video Prompting Guide, read in a browser 17 Sep 2026. An earlier version of this paragraph said AI video models weight earlier tokens heavily, with no source.

✕ UNSTRUCTURED
A woman in an ivory blazer walks through a golden hour boulevard. Camera tracking. Warm natural light. 8 seconds.
✓ BRACKET CHECKLIST
[MOTION]: slow graceful walk, natural heel-to-toe gait [SUBJECT]: 1woman, ivory blazer, consistent face [ENV]: golden hour boulevard, warm backlight [CAMERA]: tracking dolly [STYLE]: warm natural light, fine grain [DURATION]: 8s
TAKEAWAY

This structure puts motion first so it is never left out. Veo 3.1 and Seedance 2.5 publish different orders, so check your vendor's guide.

Identity Lock — The Most Important Line in Any Video Prompt

Temporal consistency — keeping the subject looking the same across every frame — is the most common failure mode in AI video. Without explicit instructions, the model will subtly change the subject's face, hair, and clothing between frames. Over 8 seconds, this drift is obvious and unusable.

Identity lock is a prompt instruction — not a model setting. You add it explicitly to the [SUBJECT] section:

[SUBJECT]: 1woman, consistent face and identity throughout every frame, same hair, same clothing, no morphing, no identity drift, same features in every frame

The phrase "no morphing" is particularly effective — it directly names the failure mode you're preventing. For models that support reference image upload (Higgsfield Studio, ControlNet-based pipelines), use a reference photo in addition to the text lock. Text alone may not be enough for high-movement sequences.

TAKEAWAY

Identity lock is a prompt line, not a setting. Add it to the subject with "no morphing", and use a reference photo too where the model supports one.

Camera Specification

The camera bracket controls how the scene is filmed — not what's in it. Separating camera instructions from subject instructions is essential. When they're mixed, the model often applies movement instructions to the subject rather than the camera.

[CAMERA]: tracking dolly, 35mm anamorphic lens, shallow depth of field

Key camera movements for AI video:

Tracking dolly — follows the subject laterally. Produces the most natural-looking motion for walking sequences. Dolly push-in — moves toward the subject. Creates intensity and intimacy. Handheld — introduces authentic shake. Good for UGC and documentary aesthetics. Aerial drone — overhead or elevated perspective. Static tripod — zero camera movement. Forces the subject's motion to carry the scene.

Adding a lens reference (35mm anamorphic, 85mm portrait, 24mm wide) changes the optical quality of the output significantly — the model understands the depth, compression, and bokeh characteristics associated with each lens.

TAKEAWAY

Keep camera instructions apart from the subject. Name the movement, such as tracking dolly or static tripod, and add a lens to shape depth and bokeh.

Duration and Frame Rate

Clip lengths differ by model. Runway: “Gen-4 creates videos in 5 or 10 second durations based on an input image and text prompt you provide.” Its newer Gen-4.5 is more flexible: “Supported durations 2 - 10 seconds”.Runway, Creating with Gen-4.5, read at source 22 Sep 2026. Kling: “The new model generates up to 15 seconds of continuous video, with a flexible duration ranging from 3 to 15 seconds.” Dreamina, for Seedance 2.5: “You can create cinematic videos up to 30 seconds in standard mode or extend them to 180 seconds with the beta long-video mode.” Always specify duration, in the prompt or the tool’s duration setting, so the motion you describe fits the clip you get.Runway, Gen-4 Video Prompting Guide, read in a browser 17 Sep 2026 · Kling AI, Kling VIDEO 3.0 model user guide, read at source 17 Sep 2026 · Dreamina (ByteDance), Seedance 2.5, read at source 17 Sep 2026. An earlier version of this paragraph said most models support clips up to 10–15 seconds and default to their shortest length, with no source.

[DURATION]: 8 seconds, 24fps, no scene cuts, continuous motion throughout

"No scene cuts" is important. Without it, some models will insert a hard cut in the middle of the clip — which destroys the continuity you were trying to build. "Continuous motion throughout" reinforces that the action should flow from start to finish without interruption.

TAKEAWAY

Always state duration. Clip lengths differ by model, so match the motion to the clip; add "no scene cuts" so some models do not insert a hard cut.

A Complete Working Example

[MOTION]: slow graceful walk, natural heel-to-toe gait rhythm, subtle sway of fabric [SUBJECT]: 1woman, consistent face and identity throughout, no morphing, ivory silk blazer, same features every frame [ENV]: golden hour boulevard, warm amber backlight from behind, long shadows across pavement, soft bokeh background [CAMERA]: tracking dolly alongside subject, 35mm anamorphic lens, horizontal lens flare from backlight [STYLE]: soft natural backlight, cinematic warm grade, fine film grain [DURATION]: 8 seconds, 24fps, no scene cuts, continuous motion

More prompts, when something changes.

Prompts, model guides and workflow notes, sent when there is something new. Free.

SUBSCRIBE FREE ↗
ABOUTMETHODVERIFYPRIVACYCONTACTINDEXAI PROMPT GENEER · EVERY ARTICLE CARRIES ITS OWN CHECKED DATE