"Make a video from text" means two different things and the tools are completely different. Sorting out which one you want takes thirty seconds and saves you from evaluating the wrong category.
Prompt to clip: you describe a shot, a model generates footage of it. This is what a video model does.
Script or article to finished video: you paste text, a tool builds a narrated video from stock footage, templates and slides. This is an assembly tool, not a generator.
TL;DR
- Prompt to clip is generation. On 8frame from 24 credits per 5 seconds
- Script to video is assembly: stock footage, templates, synthetic narration. Different tools entirely
- Free tiers in both watermark, cap length and limit resolution
- The hybrid most people actually want: write the script, generate the visuals, narrate it yourself or with TTS
- On 8frame: we generate clips, not assembled videos. There is no script-to-video pipeline here
Which one do you want
You want prompt to clip if you have a specific shot in mind: an atmospheric plate, a product in motion, a visual you cannot film. You are making footage.
You want script to video if you have an article, a blog post or a script and want a narrated video with visuals over it, quickly, without making creative decisions per shot. You are making a package.
The second category is what most "turn text into video" marketing means, and the output is recognisable: stock footage, a synthetic voice, simple captions, competent and generic.
What free tiers restrict
Both categories: a watermark, which rules out client or ad delivery. Short duration or output length caps. Reduced resolution. Slow queues, which matters more than people expect since iteration count drives quality.
Script to video specifically: free tiers usually limit you to a small stock library and the platform's own voices, and export length is capped low.
Commercial use is frequently not granted on free tiers in either category. Worth reading before you build anything on one.
The genuinely unrestricted route is self-hosting an open-weights model: Wan 2.2 under Apache 2.0, or LTX-2.5 under a community licence that permits commercial use at no cost under $10M annual revenue. No watermark, no cap, your own GPU.
What generation costs when free runs out
Canon 8frame prices, a credit is $0.01 at pack rate:
| Model | Clip | Credits | Per second |
|---|---|---|---|
| Wan 3.0 at 480p | 5s | 24 | 4.80 |
| Veo 3.1 Lite, silent | 8s | 33 | 4.13 |
| Wan 3.0 Prime at 480p | 5s | 46 | 9.20 |
| Veo 3.1 Lite, with audio | 8s | 54 | 6.75 |
| Kling v3 Standard, silent | 5s | 57 | 11.40 |
| Kling v3 Standard, with audio | 5s | 86 | 17.20 |
| Gemini Omni Flash, with audio | 8s | 135 | 16.90 |
| Veo 3.1 Standard | 8s with audio | 448 | 56.00 |
Veo 3.1 Lite at 6.75 credits per second with audio is the cheapest good option by a wide margin, and 4.13 silent is the cheapest anything.
The hybrid most people actually want
Neither category alone, but a combination that costs a few dollars:
- Write the script yourself, to a word count. 150 words per minute of finished video, so 230 words for 90 seconds. This is the part that determines whether anyone understands or stays.
- Narrate it. ElevenLabs TTS at 13 credits per 1,000 characters, so a 90-second script is about 18 credits. Or record yourself, which is free and better when the content is about your judgement.
- Generate the visuals per beat. Illustrative stills at 6 credits on Seedream 5, with motion added in the edit, or plates at 24 credits on Wan 3.0 at 480p.
- Assemble. Merge Videos at 10 credits per run, or any editor.
- Caption it, because most views are muted.
A 90-second video this way: narration (18), twelve stills (72), two plates (48), a music bed (13), assembly (10) is about 161 credits, roughly $1.61.
That is cheaper than most script-to-video subscriptions and the visuals are specific to you rather than drawn from the same stock library everyone else uses.
Where we fit, plainly
8frame generates clips, images and audio. It does not turn a script into a finished video. There is no article-to-video pipeline, no template library, and no automatic narration-plus-stock assembly. Merge Videos at 10 credits combines clips with ffmpeg transitions, which is basic assembly rather than a production tool.
If what you want is paste-text-get-video, an assembly tool is the right category and we are not it.
The thing neither category fixes
A script with nothing in it. Production polish is cheap now, which means it is not a differentiator, and a generic narrated video over stock footage is the most abundant thing on the internet.
The scarce input is a specific point of view, and it is the part no tool supplies.
FAQ
What is the best free text to video generator? For prompt-to-clip, free tiers of the major models to evaluate quality. For script-to-video, an assembly tool. They are different categories.
Does 8frame turn a script into a video? No. It generates clips, images and audio. Assembly is basic.
What is the cheapest way to make a narrated video? Write the script, TTS at 13 credits per 1,000 characters, stills at 6 credits with motion in the edit. About $1.61 for 90 seconds.
Can I get watermark-free output for free? Only by self-hosting an open-weights model on your own GPU.
Write the script, generate the beats, narrate it. The 8frame canvas is free and unlimited, and generation is paid from $19/month.