A faceless video is a script, a voice, and a visual layer that holds attention while the voice does the work. All three parts are now cheap, and that is both the opportunity and the problem: the format's barrier to entry collapsed, so the only remaining differentiator is whether the script is worth listening to.
Here is the full stack with real costs, and an honest account of where the format tops out.
TL;DR
- Script first, always. It is the entire product. Visuals are pacing, not content
- Voice: ElevenLabs TTS at 13 credits per 1,000 characters, so a 10-minute script runs about 130 to 195 credits
- Visuals: Wan 3.0 at 480p, 24 credits per 5 seconds, is the workhorse for cutaways
- A full 10-minute video: roughly 700 to 900 credits, about $7 to $9
- The ceiling: faceless channels compete on volume in a category where volume is now free, which is a hard position
The stack, part by part
1. Script
A 10-minute video is roughly 1,300 to 1,500 spoken words. This is where all the value sits and it is the part no tool shortcuts: a well-made faceless video with a boring script is a boring video with good pacing.
The structural rule that matters: one idea per 60 to 90 seconds, and state the payoff early. Retention on faceless content falls off hardest in the first 30 seconds, and the cause is almost always a slow open rather than weak visuals.
2. Voice
ElevenLabs TTS on 8frame is 13 credits per 1,000 characters. A 1,400-word script is roughly 8,000 characters, so about 104 credits, a bit over a dollar.
What to get right:
- Pick one voice and keep it. The voice is the channel's identity in this format. Switching voices between videos resets recognition.
- Punctuate for breath. Synthetic delivery follows punctuation more literally than a human does. Commas and full stops are your pacing controls.
- Write numbers as words where you want them spoken a specific way. "$19" is ambiguous; "nineteen dollars" is not.
- Listen to the whole track once. Names and jargon are where synthetic voices mispronounce, and those are the words that matter.
3. Visuals
The visual layer's job is to prevent the eye from leaving, not to explain. Three sources, in order of how much you should use them:
Screen recordings and real footage. Free, specific, and more trustworthy than anything generated. If the video is about a tool, a process or data, show the actual thing.
Generated cutaways. Wan 3.0 at 480p is 24 credits per 5 seconds. For a 10-minute video needing 30 cutaways, that is 720 credits. Trim it by reusing plates and holding some shots longer.
Stills with motion. Cheaper than video: generate an image at 6 credits on Seedream 5 and add a slow push in the edit. Twenty stills cost 120 credits and fill a lot of runtime.
4. Music and sound
A bed under the whole thing plus a few transitions. ElevenLabs Music is 81 credits per minute, so 10 minutes of original bed is 810 credits, which is too much: use a short loopable section instead, or a music library. ElevenLabs SFX at 5 credits each covers transitions.
What a 10-minute video costs
Canon 8frame prices, a credit is $0.01 at pack rate:
| Element | Route | Credits |
|---|---|---|
| Voiceover, 8,000 characters | ElevenLabs TTS at 13 per 1k | 104 |
| 20 generated stills with motion added in edit | Seedream 5 at 6 | 120 |
| 15 generated cutaways, 5s each | Wan 3.0 480p at 24 | 360 |
| 2 minutes of music bed, looped | ElevenLabs Music at 81/min | 162 |
| 6 transition effects | ElevenLabs SFX at 5 | 30 |
| Thumbnail, 12 candidates | Seedream 5 at 6 | 72 |
| Total | ~848 credits, about $8.48 |
So one 10-minute faceless video is most of a $19 Starter month. At a weekly cadence the $49 Creator plan's 3,000 credits is the honest tier, and you would lean harder on reused plates and stills.
The cheaper version that works just as well
Cut the generated video entirely:
| Element | Route | Credits |
|---|---|---|
| Voiceover, 8,000 characters | ElevenLabs TTS | 104 |
| 30 stills, motion added in the edit | Seedream 5 at 6 | 180 |
| Thumbnail candidates | Seedream 5 at 6 | 72 |
| Total | ~356 credits, about $3.56 |
Stills with a slow push and good cuts are visually indistinguishable from generated video in a 10-minute talking piece, because nobody is studying the background. This is the version to start with.
The honest ceiling
You are competing on volume in a category where volume is now free. Every part of this stack is available to everyone at the same price, which means the format's output is expanding faster than attention. Channels that do well in it either have a genuinely specialised knowledge base, or they are early in a niche, or they are spending on distribution.
No parasocial asset. The durable thing a channel builds is a relationship with a person. Faceless formats build a relationship with a topic, which is weaker and more easily substituted. A viewer who likes your topic will watch a competitor's video about the same topic without hesitation.
Platform policy is a live risk. Guidance on mass-produced and low-effort content has been tightening across platforms. Faceless is not the problem; interchangeable is. Read your platform's current monetisation guidance directly rather than trusting any article's summary of it, including this one.
The realistic reason to do it anyway: you have real expertise and genuinely do not want to be on camera. That is a good reason, and in that case the specialisation carries the channel and the format is fine.
FAQ
What is the cheapest faceless setup? Voiceover plus generated stills with motion added in the edit: about 356 credits, $3.56, for a 10-minute video.
Do faceless channels still get monetised? Policies target low-effort and mass-produced content rather than the absence of a face. Check your platform's current guidance.
Should I use an AI avatar instead of no face? If the alternative is faceless, an avatar adds a synthetic presenter without adding trust. Either film yourself or stay faceless; the middle option gets the costs of both.
Write the script, then spend a few dollars on everything else. The 8frame canvas is free and unlimited, and generation is paid from $19/month.