A generation prompt is a shot brief, and it works best written in the order a cinematographer would think: subject, framing, light, motion, duration, treatment. Prompts written as adjective piles ("stunning cinematic 4k beautiful") underperform prompts written as specifications, because adjectives describe a feeling and specifications remove defaults.
The six clauses
1. Subject and action. What is in frame and what happens. One action, not three. "Hand lifts a ceramic mug from a wooden table."
2. Framing. Wide, medium, close, extreme close. Plus the angle if it matters.
3. Light: source, direction, quality. The highest-leverage clause and the most frequently omitted. "Single hard window source from camera left, deep unfilled shadow on the right." Name what stays dark.
4. Camera motion. Static, push, pan, handheld, and how fast. Include the imperfection if you want it to look filmed: "slight drift, settles at the end."
5. Duration. Models bill and behave per second. State it.
6. Treatment. Lens feel, grade direction, film or digital character. One or two clauses, not a paragraph.
The template
[Framing] of [subject] [doing one action], [setting]. [Light source], [direction], [quality], [what falls into shadow]. [Camera motion, with imperfection]. [Duration]. [Lens and treatment].
Worked:
Medium close-up of hands folding a linen napkin, restaurant table. Single overhead source, slightly behind, hard, edges of the fabric catch light and the table falls dark. Static with a small handheld drift. 5 seconds. 50mm feel, slightly warm.
What fails in every model
Compound sequences. "She picks up the mug, walks to the window, then turns around." Models handle a single continuous action far better than a sequence of state changes. Split it into shots.
Negations. "No text, no people, no logos" frequently produces the thing named. Describe what should be there instead. Where the model supports a dedicated negative field, use that rather than the main prompt. See what is negative prompting.
Precise counts. "Exactly three apples" produces two or four often enough to be unusable as a requirement.
Text in the image. Improving, still unreliable. Add text in the edit.
Iterating without starting over
Change one clause at a time. A prompt that half-works and gets four simultaneous edits cannot be debugged, and you lose the version that was nearly right.
Keep the prompts that worked. For generated assets the prompt plus the references is the source file: without them you cannot make version two. See how to organize creative assets.
When to use a reference instead of words
If identity, an exact product, or a specific look must hold, a reference image does in one input what a paragraph of description does badly. Reference-conditioned generation is the difference between "a ceramic mug" and "this mug." Models on the canvas take one to seven references depending on the family.
FAQ
How long should a video prompt be? As long as it takes to cover the six clauses, usually two or three sentences. Longer is only better when the extra words remove a default rather than add an adjective.
Should I copy prompt templates from other people? As structure, yes. As content, no: the specifics are what make a prompt work and they are specific to your shot.
Why does the same prompt give different results? Generation is stochastic. That is why testing means three runs, not one. See how to test an AI video model.
Specify, do not adjectivise. The 8frame canvas is free and unlimited, and generation is paid from $19/month.