A video model does not read your prompt as a list of instructions to satisfy. It conditions a generation on your text, weighting some parts heavily and others barely at all. When the output ignores something you asked for, it is usually one of six identifiable reasons rather than randomness.
TL;DR
- Too many instructions. Adherence drops sharply past a handful of requirements
- Negatives get dropped. "No people" frequently produces people
- Vague language has no purchase. "Cinematic" steers almost nothing
- The model tier. Adherence is a capability you pay for
- Compound motion. One camera move per clip, always
- Contradictions you did not notice, which the model resolves by picking one
1. Too many requirements
Every element in a prompt competes for influence. A prompt with twelve requirements will land maybe six of them, and which six is not predictable.
The fix: rank your requirements and put the non-negotiable ones first. Generate, check whether the top three landed, then add one more requirement per iteration rather than writing everything at once.
2. Negatives frequently invert
"No text", "without people", "no camera movement" often produce exactly the thing you excluded. The model conditions on the concepts present in the text, and mentioning something makes it more likely, not less.
The fix: describe what should be there instead. Not "no people" but "an empty room". Not "no camera movement" but "locked-off static camera". Positive description of the desired state works; negation is unreliable.
Some models offer a separate negative prompt field, which is more reliable than negation inside the main prompt because it is handled differently.
3. Vague adjectives steer nothing
"Cinematic", "beautiful", "high quality", "professional", "4K" and "masterpiece" are nearly inert. They appear in so much training caption text, attached to so many different images, that they carry almost no specific signal.
The fix: replace every vague adjective with a concrete one.
| Vague | Concrete |
|---|---|
| Cinematic | 85mm, shallow depth of field, hard backlight |
| Beautiful lighting | Single hard light from camera left, long shadows |
| High quality | Nothing. Delete it |
| Dynamic | Slow dolly in |
| Professional | Delete it |
Six concrete specifics out-steer twenty adjectives.
4. Model tier
Prompt adherence is a real capability difference and it costs money. Canon 8frame prices, a credit is $0.01 at pack rate:
| Model | Clip | Credits | Adherence |
|---|---|---|---|
| Wan 3.0 at 480p | 5s | 24 | Loose. For blocking |
| Veo 3.1 Lite | 8s silent | 33 | Good for the price |
| Kling v3 Standard | 5s silent | 57 | Reliable on simple direction |
| Kling v3 Pro | 5s with audio | 114 | Tighter |
| Seedance 2.5 | 5s at 720p | 157 | Strong, especially physical detail |
| Veo 3.1 Standard | 8s with audio | 448 | Best adherence available |
If a prompt lands on Veo 3.1 Standard and not on Wan 3.0, that is not a prompt problem.
5. Compound camera moves
Asking for a dolly in while panning right produces incoherent motion on every model at every price. This is the most reliable way to get output that ignores your direction.
The fix: one move per clip. If you need a compound move, generate two clips and cut, which is what a real edit does anyway.
6. Contradictions
Prompts frequently contain conflicts the writer did not notice: "wide establishing shot" and "close detail on the label", "bright airy daylight" and "moody shadows", "static camera" and "we follow her as she walks".
The model resolves the conflict by picking one, and which one looks like it ignored you.
The fix: read your prompt back asking only "can all of these be true in one shot?"
The order that works
- Shot size. "Medium close."
- Subject and one action. "A hand lifts a ceramic cup."
- One camera move. "Slow push in."
- Lens. "85mm, shallow depth of field."
- Light, named by source and direction. "Hard morning sun from camera left."
- What stays still. "Static background, no other motion."
Six lines, concrete, no adjectives doing nothing. Then iterate by changing one at a time so you learn which change did what.
The cheapest way to debug a prompt
Draft on Wan 3.0 at 480p for 24 credits. Three attempts is 72 credits, under a dollar, and it tells you whether the composition and the direction are landing before you spend 157 or 448 on a finish.
Debugging at 448 credits a try is the most expensive habit in this category.
FAQ
Why does my prompt work on one model and not another? Adherence is a capability difference between tiers. It is also why blocking cheaply then finishing expensively works.
Do negative prompts work? A dedicated negative prompt field is more reliable than negation inside the main prompt, which frequently inverts.
How long should a video prompt be? Long enough for six concrete specifics. Adding more requirements past that reduces how many land.
Does "4K" or "masterpiece" in the prompt help? No. Delete them.
Six concrete lines, one move, debug at 24 credits. The 8frame canvas is free and unlimited, and generation is paid from $19/month.