One learning objective per video. That single rule fixes most training video problems, because the failure mode of the format is a twenty-minute film covering nine things where the viewer retains two. Board each objective as its own short piece, and a course becomes a set of findable, rewatchable units instead of a monolith nobody scrubs through.
The structure per objective
1. Why this matters, in one sentence. Not a preamble. The consequence of not knowing it.
2. Demonstrate first. Show the whole task at normal speed before explaining any step. Learners need the shape before the detail, and explaining first produces viewers who follow along without understanding what they are building toward.
3. Break it into steps. Three to six, each with a visible state change so the learner knows a step completed.
4. Show the common mistake. The most valuable thirty seconds in most training videos and the most commonly omitted.
5. Recap the shape, not the steps. A summary that repeats all six steps is a second video.
Three to six minutes per objective. Beyond that, split.
Chunking rules that hold
One idea per screen. If the frame shows two things, the viewer picks one.
Pause after each state change. A beat of nothing is where comprehension happens; continuous narration prevents it.
Repeat the objective at the end. Not the content, the objective.
Make it scrubbable. Chapter markers, on-screen step numbers, and a title card per step, so someone returning to find step four can find step four.
What to generate
Training content has a specific risk: it is treated as authoritative, and errors in it propagate into how people do their jobs.
Generate: context and environment shots, abstract illustrations of concepts, transitions, and title backgrounds. Narration is a good fit here too, since training scripts are long and revised often: ElevenLabs TTS at 13 credits per 1,000 characters means a five-minute narration is a few cents and a script change is a re-render rather than a re-record.
Do not generate: the procedure itself, the interface, safety-critical demonstrations, or anything a learner will copy exactly. Record real screens and real hands. A generated depiction of a procedure that is subtly wrong is worse than no video.
Say what is generated where a learner might otherwise take a generated illustration for documentation.
FAQ
How long should a training video be? Three to six minutes per objective. Longer material becomes a series, which is also easier to update.
Should training videos have a presenter on camera? For soft skills and context, yes, it helps. For procedural content, the screen or the hands are the subject and a talking head competes with them.
Can AI narrate training content? Yes, and it is one of the better fits, because scripts change and re-recording a human is expensive. Have a person review every language before it ships.
One objective, demonstrated before explained, with the common mistake shown. The 8frame canvas is free and unlimited, and generation is paid from $19/month.