Lip sync is the one AI video capability that is genuinely better than the thing it replaces: redubbing a video into another language used to mean re-shooting or accepting visible mismatch, and now it does not. Free tiers exist, restricted the same way avatar tiers are, and the quality difference between tools shows up in exactly two places: consonants and the edges of the mouth.
TL;DR
- Free tiers: typically watermarked and capped around 30 seconds
- What separates tools: plosive consonants (p, b, m) and the boundary between mouth and face. Everything else looks fine everywhere
- Strongest use case: localisation of an existing video, not making a new one
- Paid calibration: OmniHuman on 8frame is priced per second of output, and a short clip plus generated voice lands around 125 credits, about $1.25
- The consent line: applying lip sync to a real person's face requires their permission, in every jurisdiction and on every serious platform
What makes sync convincing
Plosives. P, b and m require the lips to fully close. Tools that cheat this produce a mouth that never quite shuts, and it reads as wrong even to viewers who cannot say why. This is the single best test of a lip sync tool: record "papa bought a map" and watch the closures.
Mouth boundary. The edge where generated mouth meets real face. Under-performing tools produce a visible rectangle or a slight colour shift.
Jaw movement. Speech moves the whole lower face, not just lips. Sync that animates only the mouth looks like a puppet.
Timing against the audio, not the transcript. Good tools follow the waveform, including pauses and breaths.
Where free tiers restrict
Watermark, duration cap (commonly 30 seconds), resolution cap, and sometimes a limit to the platform's own stock faces rather than your footage. Test with your actual footage before committing: sync quality varies a lot with the source, especially with head movement and partial occlusion.
The strongest use case: localisation
Dubbing an existing video into other languages is where this earns its place. The speaker is real, the performance is real, and only the mouth changes to match a translated track. That is a modest, honest intervention and it replaces a genuinely expensive alternative.
The rule that keeps it honest: have a native speaker listen to every language before it ships. Synthetic voices are fluent-sounding in dozens of languages, and fluent-sounding is not correct: emphasis, register and idiom go wrong in ways the tool cannot flag. See how to localize creative for multiple markets.
The consent line
Applying lip sync to a real person means putting words in their mouth, literally. It requires that person's permission, and the permission should cover the specific words, not just the technique. A performer who agreed to a product video did not agree to a different script in a different language unless that was stated.
This is the capability most likely to be misused and the one where the rules matter most. Never apply it to a public figure, never to footage you did not have rights to, and never to change what someone said.
FAQ
Is there a free lip sync tool with no watermark? Rare. Free tiers watermark because the output is immediately usable.
Can lip sync fix bad audio? No. It changes the picture to match audio; the audio has to be right first.
What is the best test of a lip sync tool? Plosive consonants. Record a line full of p, b and m sounds and watch whether the lips actually close.
Localise real performances, get a native speaker to listen, and never change what someone said. The 8frame canvas is free and unlimited, and generation is paid from $19/month.