On 8frame, eight video models accept an audio track you bring (a voiceover, a music bed, a line of dialogue), and they do four different things with it. Pruna p-video and p-video Draft build the clip around the track: the clip runs as long as the track, up to 20 seconds, and you pay for the track's length. Creatify Boreal keeps your track in place of its generated speech, but uses only as many seconds of it as the length setting, which starts at 5. Seedance 2.5 and 2.0 take it as a reference for the generation and switch their own sound off. Pruna Avatar, Hailuo H3 Max Lip Sync and OmniHuman 1.5 animate a face to it. Every other video model, including Veo 3.1, every Kling, Wan, Gemini Omni Flash, Grok Imagine and the base Hailuo H3, has no audio input; for those, lay your track under the finished clip in the Video Editor. Everything below was read from 8frame's own code on 2026-10-01.
TL;DR
- The track sets the length: Pruna p-video and Draft. 1 to 20 seconds, billed by the track: a 10-second track is 27 credits at 720p, 7 on Draft. A longer track is refused before you are charged
- The length setting sets the length: Creatify Boreal uses the first N seconds of your track, where N is the length setting (default 5). Set it to 10 for a 10-second track: 14 credits at 720p
- The track guides, the model goes quiet: Seedance 2.5 and 2.0 take one reference track on the canvas and switch their own sound off. The price follows the length setting, not the audio: 313 for 10 seconds at 720p on 2.5
- The track drives a face: Pruna Avatar (34 for 10 seconds at 720p), Hailuo H3 Max Lip Sync (108 at its default 768p), OmniHuman 1.5 (224)
- No audio input: every other video model. Add your track afterwards in the Video Editor
- Formats: MP3 or WAV on every model that takes a track; p-video also accepts FLAC
Every video model and your audio
Prices are credits for a 10-second track at the resolutions shown; a credit is about $0.01 at pack rate. "Not accepted" means the node shows no audio input for that model.
| Model on 8frame | Accepts your audio? | What it does with it | Length rule | 10 s track |
|---|---|---|---|---|
| Pruna p-video | Yes, optional | Conditions the video on the track | Clip = the track, rounded up to the second, 1 to 20 s; longer refused before charging | 27 / 54 (720p / 1080p) |
| Pruna p-video Draft | Yes, optional | Same as p-video, in draft mode | Same as p-video | 7 / 14 (720p / 1080p) |
| Pruna p-video 2 Pro (Speed, Quality) | Not accepted | The input is hidden; a track is refused | n/a | |
| Creatify Boreal | Yes, optional | Kept in the output instead of generated speech (per the provider) | Clip = the length setting, 1 to 20 s, default 5; the first that-many seconds of the track are used, a shorter track is padded with silence; tracks up to 60 s accepted | 14 / 41 / 162 (720p / 1080p / 2K), with the length set to 10 |
| Seedance 2.5 | Yes, optional, one track | Reference for the generation; Seedance's own sound is switched off | Clip = the length setting, 4 to 30 s; a track over 30 s is cut to its first 30 | 139 / 313 (480p / 720p); 1,307 at 720p with a reference video |
| Seedance 2.0 | Yes, optional, one track | Same as 2.5 | Clip = the length setting, 4 to 15 s; a track over 15 s is cut to its first 15 | 96 / 216 (480p / 720p); 264 at 720p with a reference video |
| Pruna Avatar | Required | Lip-syncs a portrait to the track | Clip = the track, rounded up, up to 35 s; longer refused before charging | 34 / 61 (720p / 1080p) |
| Hailuo H3 Max Lip Sync | Required, from your 8frame library | Lip-syncs a still to the track | Clip = the track, 5 s minimum (shorter refused); the provider clips it at 14.8 s | 68 / 108 / 216 / 432 (480p / 768p / 1080p / 2K) |
| OmniHuman 1.5 | Required | Animates an image to the track | Up to 35 s (longer refused), billed in 5-second steps | 224 |
| Veo 3.1 (Lite, Fast, Standard), Kling v3, O3, 2.6 Pro and older, Wan 3.0, 3.0 Prime, 2.5, 2.2, Hailuo H3, Max Turbo, Max Camera Controls, Gemini Omni Flash, Grok Imagine, Happy Horse, Runway Gen-4 Turbo, Runway Act Two, ID-V2V | Not accepted | n/a |
Two of those "not accepted" rows have a near miss. Wan 3.0 Prime's provider API has a reference-audio field, but 8frame does not send it. Runway Act Two is driven by a performance video, not an audio track.
Pruna and Creatify Boreal are new to 8frame, and none of the lip sync tools had production runs in the 30 days to 2026-09-30, so we quote prices and rules here, not quality.
The track sets the length: Pruna p-video
p-video and p-video Draft are the only models on 8frame where your track decides how long the clip is. Pruna's schema describes the input as "Input audio to condition video generation" and says the duration setting is "Ignored when audio is provided". 8frame follows that: with a track attached it does not send a length at all, and bills the track instead.
- Billing: 8frame measures the track itself and bills it rounded up to the whole second. A track it cannot measure bills the 20-second ceiling: 54 credits on p-video at 720p.
- Limit: 20 seconds. A longer track is refused before you are charged, not cut short, so trim it first.
- Prompt: still required. A start image is optional.
- p-video 2 Pro: has no audio field. The node hides the input for that variant and the backend refuses a track.
Pruna's own page lists music videos as a use ("combine your own audio with generated visuals"). We have not checked an 8frame output with a track attached, so whether the file plays your track unchanged is not confirmed. Settings and the sound question on Draft are in how to use Pruna p-video.
The length setting sets the length: Creatify Boreal
Boreal works the other way round. Per fal's API page, your track "is preserved in the output instead of generated speech", and "the first duration seconds of the audio are used; shorter audio is padded with silence." The clip is as long as the length setting, whatever the track.
8frame always sends that setting, and it starts at 5 seconds, so a 10-second voiceover attached without changing it comes back as its first 5 seconds. Set the length to the track (up to 20 seconds) before generating. Boreal bills the length setting, not the track: 10 seconds is 14 credits at 720p, 41 at 1080p, 162 at 2K. 8frame accepts tracks up to 60 seconds for Boreal, but only the first 20 can ever be used. The prompt tags and the rest of the settings are in how to use Creatify Boreal.
The track guides, the model goes quiet: Seedance
Seedance 2.5 and 2.0 take your track as a reference. ByteDance's schema on Replicate describes reference audio as "for audio-driven generation and lip-sync", and tells you to refer to it in the prompt as [Audio1]; the canvas node sends one track, so that is the only tag you need. Three rules come from 8frame's code and that schema:
- Seedance's own sound goes off. Whenever a reference track is attached, 8frame tells Seedance not to generate audio, whatever the node's sound switch says. Whether your track itself is carried into the output file is not confirmed; we have not checked one.
- The length setting still decides the clip. 8frame does not change it to match the track, so set it yourself. A track longer than the version allows (30 seconds on 2.5, 15 on 2.0) is cut to its first 30 or 15 seconds before it is sent.
- Audio alone is not enough. Per the provider's schema, a reference track needs at least one reference image or a reference video with it, and on 2.5 it cannot be combined with an end frame.
The audio does not change the price. A reference video does: it moves Seedance to its video-input rate, 1,307 credits for 10 seconds at 720p on 2.5 against 313 with reference images. So the cheaper way to give Seedance your track is with a reference image, not a reference video.
The track drives a face: lip sync tools
Pruna Avatar, Hailuo H3 Max Lip Sync and OmniHuman 1.5 need a portrait and a track, and the output follows the track. They differ on limits and on how they round:
- Pruna Avatar: rounds up to the second, up to 35 seconds; longer is refused before charging. A track 8frame cannot measure bills the 35-second ceiling (119 at 720p).
- Hailuo H3 Max Lip Sync: takes no prompt. The track must be a file in your 8frame library (upload it, or pick it there). Under 5 seconds is refused; over 14.8 the provider cuts it, and the charge stops at 15 seconds.
- OmniHuman 1.5: bills in 5-second steps, so a 10.4-second track bills as 15 seconds, 336 credits.
The full price comparison, including a costed 30-second talking head, is in AI lip sync price by model.
Formats and file limits
| Model | Formats | File size | Longest track |
|---|---|---|---|
| Pruna p-video, Draft | FLAC, MP3, WAV | 10 MB | 20 s (longer refused) |
| Creatify Boreal | MP3, WAV | 10 MB | 60 s accepted, first 1 to 20 s used |
| Seedance 2.5 / 2.0 | MP3, WAV | 12 MB | any; cut to 30 s / 15 s |
| Pruna Avatar | MP3, WAV | 10 MB | 35 s (longer refused) |
| Hailuo H3 Max Lip Sync | MP3, WAV | not confirmed | 5 s minimum; clipped at 14.8 s |
| OmniHuman 1.5 | MP3, WAV | 10 MB | 35 s (longer refused) |
The canvas uploads audio as MP3 or WAV, so FLAC is the one format here you would only reach on p-video. One track per node: the audio input holds a single file on every model above.
If your sound is already on a video
A few edit modes start from a clip rather than an audio file, and can keep that clip's sound: Kling 2.6 Pro Motion Control keeps the motion video's original sound by default, Kling O1 Video has a keep-audio switch (off by default), Happy Horse's video edit has a keep-original-sound switch, and on 8frame Wan 2.7 VideoEdit always asks to keep the input clip's sound. They keep the sound you brought; they do not take a separate track. What each costs is in which AI video models work from text only.
If you want generated sound instead
If you do not have a track, most of the catalog makes its own. Veo 3.1 and Kling v3 and 2.6 Pro charge for sound and let you switch it off, Seedance and Wan 3.0 Prime include it at the same price, and Grok Imagine, Gemini Omni Flash, Creatify Boreal and Pruna's p-video models always add it. Which model does what, and what sound costs on each, is in AI video models with audio, by model. If a clip came back silent, why does my AI video have no sound walks through the causes.
You can also make the track on 8frame and then use it as your own: ElevenLabs text to speech is 13 credits per started 1,000 characters, a sound effect is 1 to 6 credits by length, and a 30-second Lyria 2 music piece is 13. Prices for every voice and music model are in AI music and voice price by model.
Putting your own audio under any clip afterwards
For the models with no audio input, and for a mix you want to control, add the sound after generation. 8frame has a separate Video Editor mode, next to the canvas, with a timeline: its layers include video, text, captions and sound. Upload your voiceover or music, put the clip and the track on the timeline, line them up and render.
This route works with every video model, and your track is never handed to a model to reinterpret. Two tips:
- Generate silent where you can. On Veo 3.1 and on Kling v3, 2.6 Pro and O3, switching sound off lowers the price, and you are replacing it anyway. Grok Imagine, Gemini Omni Flash, Creatify Boreal and Pruna's p-video models always add sound, so their own audio arrives in the clip too.
- The canvas has no timeline. Its Montage node joins 2 to 10 clips with transitions for 10 credits, but takes no audio track; the Video Editor is where sound goes. Clip length limits for planning the cut are in AI video clip length limits by model.
FAQ
Can I use my own voiceover in an AI video? On 8frame, yes, on eight models: Pruna p-video and Draft, Creatify Boreal, Seedance 2.5 and 2.0, Pruna Avatar, Hailuo H3 Max Lip Sync and OmniHuman 1.5. On any other model, add the voiceover afterwards in the Video Editor.
Which AI video model makes the clip as long as my audio? Pruna p-video and p-video Draft, up to 20 seconds, and the lip sync tools (Pruna Avatar and OmniHuman up to 35 seconds, Hailuo H3 Max Lip Sync up to about 15). On Creatify Boreal and Seedance the length setting decides the clip.
Why was my voiceover cut off? On Creatify Boreal the length setting decides how much of the track is used, and it starts at 5 seconds. On Seedance a track over 30 seconds (2.5) or 15 seconds (2.0) is cut, and Hailuo H3 Max Lip Sync stops at 14.8 seconds. More causes are in why is my AI video shorter than I asked.
Can Veo 3.1 or Kling use my audio file? No. Neither has an audio input on 8frame; both generate their own sound. Generate the clip, then add your track in the Video Editor.
What is the cheapest way to make a clip to my own track? Pruna p-video Draft: a 10-second track is 7 credits at 720p. Full p-video is 27, and Creatify Boreal with the length set to 10 is 14.
Does adding my audio cost extra? Not as a surcharge. On p-video and the lip sync tools you pay for the track's length; on Creatify Boreal and Seedance you pay for the length setting, with or without a track.
Can I attach more than one track? Not on the canvas: each node takes one. Seedance's provider allows several reference tracks, but the 8frame node sends one.
Can I use my own audio through Claude or Cursor? No. The 8frame MCP connector has no audio input on any tool, so the lip sync tools and audio-driven modes (p-video with a track, Boreal or Seedance with your audio) are not available through it. Setup is in generate AI video from Claude.
Sources
- Which video nodes show an audio input, per model, and what they send: 8frame front tool registry and model compatibility map, read 2026-10-01
- What each model does with a track, formats, file sizes, length limits, trimming and pre-charge refusals: 8frame generation code for Pruna Video, Creatify Boreal, Seedance, Hailuo H3, Pruna Avatar, OmniHuman and Wan, the shared audio validators and the audio trim step, read 2026-10-01
- Prices: credit cost functions in the 8frame tool registry, read 2026-10-01
- MCP connector inputs: 8frame MCP tool definitions, read 2026-10-01
- Video Editor mode and its sound layers: 8frame front code, read 2026-10-01
- Pruna p-video audio input and duration rule: the input schema and model page of prunaai/p-video, checked 2026-10-01
- Creatify Boreal audio rule: fal.ai/models/creatify/boreal/api, checked 2026-10-01
- Seedance reference audio rules: the input schemas of bytedance/seedance-2.5 and bytedance/seedance-2.0, checked 2026-10-01
Try it without signing up: the 8frame canvas opens on a real board. The canvas is free and unlimited; generation is paid from $19/month.