Making a still image speak your own audio track costs, on 8frame, from 17 credits for 5 seconds (Pruna Avatar at 720p) to 112 (OmniHuman 1.5), with Hailuo H3 Max Lip Sync in between at 34 to 216 depending on resolution. A credit is $0.01 at pack rate. These tools take a portrait and an audio track, and the output is as long as the audio. Runway Act Two is the other route: a real performance on video drives a character, at 8 credits a second. Prices are read from 8frame's pricing code on 2026-09-30.
Lip sync tools on 8frame, one table
| Tool | You give it | Length | 5 seconds | Credits per second |
|---|---|---|---|---|
| Pruna Avatar | a portrait and an audio track | follows the audio, up to 35 s | 17 at 720p, 31 at 1080p | 3.4 / 6.1 |
| Hailuo H3 Max Lip Sync | a still and an audio track | follows the audio, 5 s minimum, up to about 15 s | 34 / 54 / 108 / 216 at 480p / 768p / 1080p / 2K | 6.75 / 10.8 / 21.6 / 43.2 |
| OmniHuman 1.5 | an image and an audio track | 5 to 35 s, billed in 5-second steps | 112 | 22.4 |
| Runway Act Two | a character (image or video) and a performance video | 3 to 30 s of reference | 40 | 8 |
None of these had production runs on 8frame in the 30 days to 2026-09-30, so we quote their prices, not their quality. Test on your own portrait and voice before committing a project.
How each one bills
Pruna Avatar bills per second of output, and the output follows your audio, measured on 8frame's side and rounded up to the next second. 720p is 3.375 credits a second, 1080p 6.075. Tracks longer than 35 seconds are refused before you're charged. The output keeps the portrait's shape.
Hailuo H3 Max Lip Sync also bills the audio's length, with a 5-second minimum, at four resolution tiers; 768p is the default. It takes about 15 seconds of audio per generation, so a longer script is split into several clips.
OmniHuman 1.5 bills in 5-second steps from 5 to 35 seconds: a 7-second track bills as 10 seconds, 224 credits.
Runway Act Two is different in kind: instead of an audio track, you give it a video of someone performing, and it transfers the performance (face and timing) to your character, at 8 credits per second of the reference.
A 30-second talking head, costed
One portrait, one 30-second script read by a synthetic voice:
| Route | Generation | Credits |
|---|---|---|
| Portrait | Nano Banana 2, 1K | 11 |
| Voice track (under 1,000 characters) | ElevenLabs text to speech | 13 |
| Pruna Avatar, 720p, 30 s | one clip | 102 |
| Pruna Avatar, 1080p, 30 s | one clip | 183 |
| Hailuo H3 Max Lip Sync, 768p | two 15-second clips | 324 |
| OmniHuman 1.5, 30 s | one clip | 672 |
So a 30-second talking head runs from 126 credits (portrait, voice and Pruna Avatar at 720p) to 696 (the same with OmniHuman). Use your own recorded voice instead of text to speech and drop the 13.
Which one to use
Cheapest, and long enough for a full script: Pruna Avatar, up to 35 seconds in one clip.
Highest resolution: Hailuo H3 Max Lip Sync at 2K (216 for 5 seconds).
When a real actor's timing matters more than a voice track: Runway Act Two, driven by a performance video.
When you don't have an audio track at all: models that generate a speaking person with sound in one pass, such as Veo 3.1 (448 credits for 8 seconds with audio), produce the model's own voice, not yours. For a specific voice, make the track first and use a lip sync tool.
Before you start
Use a portrait you have the right to animate, and a voice you have the right to use. A real person's face or voice needs their documented consent, and platforms have rules on labelling realistic synthetic people.
What to check: mouth shapes on hard consonants, teeth that stay stable, blinks that look natural, and the head not drifting away from the portrait's framing over a long clip.
FAQ
What is the cheapest AI lip sync? On 8frame, Pruna Avatar at 720p: 17 credits for 5 seconds, 102 for 30 seconds.
How much does OmniHuman cost? On 8frame, 22.4 credits a second, billed in 5-second steps from 5 to 35 seconds: 112 for 5 seconds, 672 for 30.
How long can a lip sync clip be? Up to 35 seconds with Pruna Avatar or OmniHuman, about 15 seconds per clip with Hailuo H3 Max Lip Sync. Longer scripts are split and joined (Merge Videos is 10 credits a join).
Is lip sync the same as a talking avatar? It's the core of one: a still of a person animated to an audio track. See best AI talking avatar generators for the wider picture.
Sources
- 8frame lip sync prices, lengths and billing rules: the credit cost functions in the 8frame tool registry, read 2026-09-30; production usage: 8frame production data, read-only, 2026-09-30
Make the voice track, pick the price point, check the mouth. The 8frame canvas is free and unlimited; generation is paid from $19/month.