There are two ways to get speech onto a moving mouth and they produce different results. Generated together: a model creates the picture and the audio in one pass, so the mouth was always moving to those words. Applied afterwards: a lip sync model modifies an existing face to match a separate audio track.
The first holds sync better. The second lets you use footage of a real person, which is the more useful capability.
TL;DR
- Veo 3.1 generates picture and audio jointly, which is why sync holds through camera moves. 448 credits for 8 seconds, 54 on the Lite tier
- Gemini Omni Flash generates audio by default at 135 credits for 8 seconds, with a 39-second median generation time on our jobs
- Kling v3 Standard at 86 credits includes audio; the 57-credit tier is silent
- OmniHuman applies sync to existing footage, which is the route for a real person
- The plosive test: record a line full of p, b and m and watch whether the lips actually close
The two routes
Joint generation. The model produces frames and audio together, so mouth shapes were formed for those sounds from the start. Sync survives camera movement and head turns because it was never a separate layer.
Best for: generated characters, short scenes, anything where the speaker does not need to be a specific real person.
Applied lip sync. You have footage of a real person and a separate audio track. A model reshapes the mouth region to match. This is what localisation runs on.
Best for: dubbing your own recorded video into other languages, which is the strongest use case in the entire category.
What each costs
Canon 8frame prices, a credit is $0.01 at pack rate:
| Route | Clip | Credits | Notes |
|---|---|---|---|
| Veo 3.1 Lite | 8s with audio | 54 | Cheapest joint generation |
| Kling v3 Standard | 5s with audio | 86 | Audio included at this tier |
| Kling v3 Pro | 5s with audio | 114 | Tighter adherence |
| Gemini Omni Flash | 8s with audio | 135 | Audio by default, fastest |
| Veo 3.1 Fast | 8s with audio | 168 | Most of the Veo quality |
| Veo 3.1 Standard | 8s with audio | 448 | Best sync and quality |
| Higgsfield Standard | 5s 720p | 95 | Presenter work with identity lock |
| OmniHuman | per second of output | varies | Applied sync to existing footage |
| ElevenLabs TTS | per 1,000 characters | 13 | The audio track for applied sync |
The plosive test
This is the single best evaluation of any lip sync, and it takes thirty seconds.
Record or generate a line dense with p, b and m sounds. "Papa bought a map." Those consonants require the lips to close completely. A tool that cheats produces a mouth that never quite shuts, and it reads as wrong even to viewers who cannot say why.
Watch the closures frame by frame. If the lips meet, the sync is real. If they hover, it is approximate and will look uncanny at length.
The other three tells
The mouth boundary. The edge where a generated or modified mouth meets the real face. Weak tools produce a faint rectangle or a slight colour shift.
Jaw movement. Speech moves the whole lower face. Sync that animates only the lips looks like a puppet.
Timing against the waveform, not the transcript. Good sync follows the actual audio, including pauses and breaths. Bad sync follows the text and drifts.
The consent rule
Applying lip sync to a real person means putting words in their mouth, literally. It requires that person's permission, and the permission must cover the specific words, not just the technique.
Someone who agreed to appear in an English product video did not agree to say different things in six languages unless that was stated. Never apply it to a public figure, never to footage you do not have rights to, and never to change what someone said.
This is the capability with the clearest misuse path in the whole category, and the rules are not optional.
The strongest use case
Localising your own recorded video. The speaker is real, the performance is real, and only the mouth changes to match a translated track. That is a modest intervention replacing a genuinely expensive alternative.
The rule that keeps it honest: have a native speaker listen to every language before it ships. Synthetic voices are fluent-sounding in dozens of languages and fluent-sounding is not correct. See best free AI video translators.
FAQ
Which model has the best lip sync? Veo 3.1 Standard for joint generation at 448 credits. Veo 3.1 Lite at 54 is the value option.
How do I test lip sync quality? The plosive test. Record a line full of p, b and m and watch whether the lips close.
Can I lip sync footage of myself? Yes, that is the legitimate route, and it is the best use of the technology.
Does Kling include audio? At 86 credits for 5 seconds, yes. The 57-credit tier is silent.
Test the plosives, get consent, localise real performances. The 8frame canvas is free and unlimited, and generation is paid from $19/month.