OmniHuman does one thing: it puts speech on a face that already exists. That makes it categorically different from the video models on this canvas, which generate a face along with everything else, and it makes its strongest use obvious. Localising your own recorded video is the case where a modest intervention replaces a genuinely expensive alternative.
It is also the tool on this canvas with the strictest requirement attached, so that comes first.
TL;DR
- Priced per second of output, so cost scales with clip length. Check it against your actual duration before a batch
- Strongest use: dubbing your own footage into other languages, keeping the real performance
- Consent is required and specific. Permission for one script does not cover a different one
- Pair it with ElevenLabs TTS at 13 credits per 1,000 characters for the translated track
- Test with the plosive check before committing to a batch: p, b and m require the lips to close
The consent requirement
OmniHuman puts words in a real person's mouth, literally. That requires the person's permission, and the permission has to cover the specific words, not just the technique.
Someone who agreed to appear in an English product video did not agree to say different things in six languages unless that was stated. A performer who was paid for one script was not paid for another.
Never a public figure. Never footage you do not have rights to. Never to change what someone said. This is the capability with the clearest misuse path on this canvas, and the rules are not negotiable rather than being cautious advice.
The workflow for localisation
- Record the original. A real person, a real performance. This is the asset OmniHuman preserves.
- Transcribe it. Free, and accurate enough now that this is not the bottleneck.
- Translate it properly. This is the step people treat as automatic and it is the one that decides whether the result is good. Machine translation gets grammar right and register wrong.
- Generate the translated speech. ElevenLabs TTS at 13 credits per 1,000 characters, so a 3-minute script at roughly 2,800 characters is about 36 credits.
- Apply OmniHuman to match the mouth to the new track.
- Have a native speaker listen to every language before it ships.
Step 6 is not optional. Synthetic voices are fluent-sounding in dozens of languages, and fluent-sounding is not correct: emphasis lands on the wrong word, register comes out formal where it should be casual, your product name is mispronounced consistently for the whole video. Thirty minutes of a native speaker's time per language is the cheapest quality insurance available.
The plosive test
Before you run a batch, test the sync quality on your actual footage, because it varies a lot with the source.
Record a line dense with p, b and m sounds. "Papa bought a map." Those consonants require the lips to close completely. Watch the closures frame by frame: if the lips meet, the sync is real; if they hover, it is approximate and will read as uncanny across a long video.
Three other tells worth checking: the boundary where the modified mouth meets the real face (weak results show a faint rectangle or colour shift), whether the jaw moves rather than just the lips, and whether timing follows the audio waveform including pauses rather than following the transcript.
What degrades the result
Head movement. A still or slowly moving head syncs better than one turning.
Partial occlusion. A hand near the face, a microphone, hair across the mouth.
Fast speech. Dense consonants blur, which is also why the plosive test matters.
Low resolution source. The mouth region is where detail matters most and there is no video upscaler on this canvas to recover it.
Extreme angles. Profile views are much harder than front or three-quarter.
Where it fits versus generating from scratch
Use OmniHuman when the speaker must be a specific real person: localising your own content, correcting one line in a finished video, or producing a message from a named executive who has consented.
Use a video model that generates audio jointly when the speaker does not need to be anyone real. Veo 3.1 Lite at 54 credits for 8 seconds with audio, or Gemini Omni Flash at 135 for 8 seconds, both generate picture and sound together so sync holds through camera movement because it was never a separate layer.
Use Higgsfield Standard at 95 credits for 5 seconds at 720p when you want a consistent presenter figure with identity lock rather than a specific real person.
FAQ
How is OmniHuman priced? Per second of output, so cost scales with clip length. Price your actual duration before committing to a batch.
What is the best use for it? Localising your own recorded video into other languages.
Do I need consent? Yes, and it must cover the specific words. Permission for one script does not extend to another.
How do I test sync quality? The plosive test. A line full of p, b and m, checked frame by frame for lip closure.
Localise real performances, get specific consent, check the plosives. The 8frame canvas is free and unlimited, and generation is paid from $19/month.