Leaderboards tell you which model wins on average prompts from average users, which is almost never your situation. Twenty minutes of structured testing on your own prompt tells you more, and the structure matters because the intuitive way to test (generate a few, look at them, pick the prettiest) measures the wrong thing.
The method
1. Write one prompt with four separately checkable clauses. A camera move, a light direction and source, a specific action, and a framing or duration. Each must be objectively present or absent.
Example: "Medium shot, hand lifting a ceramic mug from a wooden table. Single hard window source from camera left, deep shadow on the right. Slow push in. 5 seconds."
2. Run it three times per model. One run tells you nothing about consistency, and consistency is what determines your real cost.
3. Score clause survival, not beauty. For each generation, count how many of the four clauses actually appear. That number is prompt adherence, and it predicts your attempt count.
4. Count attempts to one usable clip. The only metric that converts to money.
5. Judge as a grid. All twelve outputs at the same size, side by side. Sequential viewing flatters whatever you saw last, which is a real and well-documented bias.
What to record
| Model | Clauses kept (of 4) | Attempts to usable | Credits per attempt | Cost per usable |
|---|---|---|---|---|
| A | ||||
| B |
The last column is the decision. A model at 86 credits landing in two attempts (172) beats a model at 54 landing in five (270), and the second one looked cheaper on the pricing page.
What not to test
Demo prompts. Vendors optimise for them, and they tell you nothing about your brief.
One spectacular output. Cherry-picking is what marketing does; your production average is what you will live with.
Quality in isolation. At the top of the market, quality clusters. Adherence and consistency separate.
Resolution. Rarely the deciding factor and frequently a distraction from motion and lighting behaviour.
The second test, once you have a favourite
Run your prompt again a week later. Models get updated silently, and a model that behaved one way in evaluation can behave differently in production. Teams that re-test quarterly catch drift; teams that do not attribute it to their own prompting.
Cost of testing properly
Four models, three runs, at a mid-tier price: roughly twelve generations at 54 to 86 credits each, so 650 to 1,000 credits, about $6.50 to $10 at pack rate. That is less than an hour of anyone's time and it settles a decision that affects every clip afterwards.
Running all four from one credit pool on one surface is what makes this practical, because the alternative is four subscriptions to compare four models once.
FAQ
How often should I re-test? Quarterly, and whenever a model family ships a new generation. The ordering changed materially more than once in 2026.
Can I trust arena rankings? As a shortlist, yes. As a decision, no: they measure average prompts and yours is not average.
How many models should I shortlist? Two or three after the four elimination questions. Testing six is a way of avoiding the decision.
Score the clauses, count the attempts, and judge the grid. The 8frame canvas is free and unlimited, and generation is paid from $19/month.