Correction, 2026-10-02: this page still compared models on things we have not measured. It called the image-to-video route "good enough most of the time", told readers to prove a shot cheaply and "finish" it on Seedance, recommended Veo 3.1 multi-reference for shots where "the face is the deliverable", set a 53-credit still-first route against 282 credits of text-to-video re-rolls as if one needed a single attempt and the other six, and said a locked still makes later shots "much easier". Those lines are now replaced with facts you can check on the canvas (how many reference images each model takes, start image and end frame support, length, resolution and price) and with the advice to compare every take against the approved still. The Seedance close-up note now matches our E005 page: in our August 2026 tests every close-up face reference failed on Seedance 2.0 and 2.5, while waist-up references with the face about a quarter of the frame height, and full-length ones, passed, and its source now names those tests instead of production runs. No price changed.
Correction, 2026-10-01: this page called Wan 3.0 reference-to-video, at 34 credits for 5 seconds at 480p, the cheapest reference route on the canvas. That was wrong. On the canvas, base Wan 3.0 animates one start image and takes no reference images, and Wan 3.0 Prime's reference mode takes a reference clip, not images. The 34-credit reference-to-video price exists only through the 8frame MCP connector, where it is also driven by a clip. The cheapest reference-image route on the canvas is Grok Imagine's reference mode, from 35 credits for 5 seconds at 480p. We also added that Seedance, Gemini Omni Flash and Happy Horse take a start image or references but not both, and that Kling v3 takes no reference images, and we turned statements about where faces drift, which we have not measured, into advice.
Correction, 2026-09-29: the first version of this page described Kling O1 Reference as a "reference pass" followed by Kling O1 Video for the clip. O1 Reference generates the clip itself; O1 Video is for editing a video you already have. It also recommended Seedance for face-locked shots without warning that Seedance refuses close-up face references, and it gave Kling 2.6 Pro's silent price without saying so. All three are fixed below.
If you're searching for an AI video consistent character setup, the answer isn't a model, it's an order of operations. Lock the character as a still image first, then drive every clip from that still. On 8frame a cheap version of that is Seedream 5 Lite at 6 credits for the portrait and Kling 2.6 Pro at 47 credits for a silent 5-second clip that opens on it: 53 credits for the first shot. Here's the full workflow, what each reference-capable video model takes on the canvas, and how to check the takes before you commit to a cut.
TL;DR
- Don't fix consistency in the video model. Fix it in the still, then animate the still.
- Cheap lock: Seedream 5 Lite at 6 credits (Seedream accepts up to 14 reference images). Nano Banana 2 costs 8 to 22 credits by resolution, takes up to 10 references and goes up to 4K.
- Reference-capable video on the canvas, cheapest first, for 5 seconds: Grok Imagine reference mode from 35 credits at 480p or 48 at 720p (1 to 10 references), Kling O1 Reference at 73 (up to 7), Gemini Omni Flash at 85 (1 to 7), Happy Horse at 95 at 720p (1 to 9), Seedance 2.5 at 157 at 720p (up to 9 on the canvas, audio included, up to 30 seconds). Veo 3.1 multi-reference is 448 for its fixed 8 seconds (up to 3 references).
- Seedance, Gemini Omni Flash and Happy Horse take a start image or references, not both. Kling v3 and Wan 3.0 take a start image and no reference images.
- In our August 2026 tests, Seedance 2.0 and 2.5 rejected every reference with a close-up face (error E005). Waist-up references, with the face about a quarter of the frame height, and full-length ones went through.
- We have not measured which model keeps a face closest, so this page does not rank them. Compare every take with the approved still, check profiles, long takes and two-character shots first, and budget for retakes.
Why text-to-video can't hold a character
Text-to-video generates a fresh face every run. Same prompt, same model, and the face can come back as a different person. For one standalone clip that doesn't matter. The moment two clips cut together, it's a production stopper. The mechanism is in what is character consistency in AI; the practical version is that a prompt is a description, and descriptions have many valid answers.
A reference image is a picture instead of a description: every shot gets the same face to work from, not a sentence with many valid answers. It does not guarantee a match, so you still compare each take with the still, but it is why the workflow below spends its first 6 credits on an image rather than a clip.
The expensive mistake is doing it in the wrong order. Re-rolling text-to-video until a face looks right costs 47 credits per silent attempt on Kling 2.6 Pro. Six attempts is 282 credits, about $2.82 at pack rate, and the face you like exists in one clip only: shot two starts from a new face again. Locking the still first costs 6 credits and gives shot two, three and twelve the same image to start from.
Step 1: lock the character as a still
Generate the canonical portrait yourself. Don't start from a stock photo, because you want a face nobody else's campaign is also using.
Seedream 5 Lite, 6 credits. Seedream accepts up to 14 reference images at once, more than any other image model on the canvas. That matters less for the first portrait and a lot for the follow-ups: once you have a face, you can feed it back alongside a wardrobe reference, a lighting reference, and a location plate. At about $0.06 an image, this is some of the cheapest identity work in the pipeline.
Nano Banana 2, 8 to 22 credits. Priced by resolution, up to 4 images per run, up to 10 references. Reach for it when you need a bigger file to crop from: it goes up to 4K, at 22 credits.
Write the prompt around what must stay fixed, not the mood. Age, build, hair length and color, eye color, skin tone, one distinguishing feature, wardrobe with a named color. Mood you can change per shot. Structure you can't.
Step 2: build an angle pack, then storyboard it
One front-facing portrait is not enough when the shots show the character from other sides. Generate the same character in the views the shots need, usually three to five: face-forward, three-quarter, profile, one from a slightly raised camera, and one waist-up or full-length shot, the framing that passed Seedance's filter in our tests. At 6 credits each on Seedream 5 Lite, a five-image pack costs 30 credits, about $0.30. That pack is the asset. Every downstream clip reads from it.
ShotGrid earns its place here too. For 42 credits it generates a 3x3 sheet of nine frames from 1 to 10 reference images, so you see your character in nine moments before spending a video credit, and you catch the shots where they fall apart at image prices instead of video prices.
Step 3: drive video from the still
Two routes. We have not measured which one keeps a face closer, so choose by what the shot needs: whether it can open on your still, how many references you want to attach, how long it runs, and the price.
Image-to-video from the locked still. Feed the portrait as the first frame and prompt the motion. Kling 2.6 Pro costs 47 credits for 5 silent seconds (95 with audio), runs 5 or 10 seconds, returns 1080p and takes no end frame; the clip opens on the face you approved. The Kling tool opens on Kling v3, which on 8frame is image-to-video only and takes no reference images (Kling O3 also needs a start image); its start frame is the only identity it gets, so make that frame from your pack too. Model-by-model breakdown in best image to video AI in 2026. Cost: 6 credits for the still plus 47 for the clip, about $0.53.
Reference-capable video models. These take several references instead of one first frame. Use them when the shot should not open on your still: the character enters later, is seen at a distance, or in profile. On Seedance, Gemini Omni Flash and Happy Horse the references replace the start image: attach both and the start image is dropped. Every cap is listed in how many reference images each model takes.
- Grok Imagine reference mode, 35 credits for 5 seconds at 480p, 48 at 720p. Takes 1 to 10 reference images; each extra image adds a fraction, so ten references cost 37 at 480p and 50 at 720p. Runs 1 to 10 seconds, takes no start image or end frame in this mode, and sound is always on. The cheapest reference route on the canvas, so a cheap place to test a pack.
- Kling O1 Reference, 73 credits for 5 seconds, 146 for 10. Generates the clip from up to 7 reference images; an attached start image is sent as another reference, not as the first frame, and there is no end frame. (Kling O1 Video, 109 credits for 5 seconds, is a different mode: it edits a video you upload.)
- Gemini Omni Flash, 85 credits for 5 seconds. Takes 1 to 7 reference images, with audio always generated. Runs 3 to 10 seconds (8 seconds is 135), returns 720p and takes no end frame.
- Happy Horse, 95 credits for 5 seconds at 720p. Takes 1 to 9 character references and runs 3 to 15 seconds, with no end frame. The tool starts at 1080p, which is 189, so set 720p while you draft.
- Seedance 2.5, 157 credits for 5 seconds at 720p (70 at 480p). Reference-driven, 4 to 30 seconds, audio on by default at the same price, so a continuous take of the character speaking doesn't need a separate pass. Up to 9 reference images on the canvas; an end frame works only with a start image and no references. The price scales with length: 30 seconds at 720p is 937. Seedance 2.0, still selectable, takes the same 9 for 48 credits at 480p or 108 at 720p, for 4 to 15 seconds. Use waist-up or full-length references, not face close-ups: in our August 2026 tests, every reference with a close-up face failed with an E005 error on both 2.0 and 2.5, retrying the same image did not help, and the same person framed from the waist up (face about a quarter of the frame height) or full-length went through. Ask for the close-up in the prompt instead. Details in Seedance flagged as sensitive (E005).
- Veo 3.1 multi-reference, 448 credits. Up to 3 reference images, 8 seconds only, 1080p, no start image or end frame in this mode, about $4.48 a generation. That is more than nine silent Kling 2.6 Pro clips (423), so price the whole sequence before you pick it.
Not on this list: Wan 3.0, which on the canvas animates one start image (Wan 3.0 Prime's reference mode takes a reference clip, not images), and Kling v3 and O3, which need a start image and take no reference images.
A realistic six-shot sequence: 30 credits for the reference pack plus six silent Kling 2.6 Pro image-to-video clips at 47 each comes to 312 credits, about $3.12. On a $19 Starter plan's 1,000 credits, which roll over while you stay subscribed, that's three sequences a month with room left over.
The comparison table
These tables show what each model takes on the canvas, not how well it keeps a face. That part you check yourself, against the approved still.
Stills
| Tool | Reference images | Credits | ~$ at pack rate | Notes |
|---|---|---|---|---|
| Seedream 5 Lite | Up to 14 | 6 | ~$0.06 | Flat price per image |
| Nano Banana 2 | Up to 10 | 8 to 22 | ~$0.08 to $0.22 | Priced by resolution, up to 4K, up to 4 images per run |
| ShotGrid | 1 to 10 | 42 | ~$0.42 | Nine frames in one 3x3 storyboard sheet |
Video
| Model | Reference images | Start image / end frame | Length | Resolution | Credits | ~$ at pack rate |
|---|---|---|---|---|---|---|
| Kling 2.6 Pro | None | Start image; no end frame | 5 or 10 s | 1080p | 47 (5 s, silent), 95 with audio | ~$0.47 |
| Grok Imagine reference mode | 1 to 10 | Neither in this mode | 1 to 10 s | 480p or 720p | 35 (5 s, 480p), 48 (720p) | ~$0.35 |
| Kling O1 Reference | Up to 7 | A start image is sent as a reference; no end frame | 5 or 10 s | Not selectable | 73 (5 s) | ~$0.73 |
| Gemini Omni Flash | 1 to 7 | Start image or references; no end frame | 3 to 10 s | 720p | 85 (5 s) | ~$0.85 |
| Happy Horse | 1 to 9 | Start image or references; no end frame | 3 to 15 s | 720p or 1080p | 95 (5 s, 720p) | ~$0.95 |
| Seedance 2.5 | Up to 9 | Start image or references; end frame only with a start image | 4 to 30 s | 480p or 720p | 157 (5 s, 720p) | ~$1.57 |
| Veo 3.1 multi-reference | Up to 3 | Neither in this mode | 8 s | 1080p | 448 (8 s) | ~$4.48 |
If the character also has to speak, the audio side is covered in AI video models with audio, by model.
What to check before you cut
No model here guarantees the same face in every clip, and we have not measured drift model by model, so treat these as checks rather than rules. The workflow is in AI character consistency workflow.
Against the approved still. Put every take next to the still it came from and check the traits you fixed in the prompt: face shape, hair length and color, eye color, skin tone, the distinguishing feature, the wardrobe color. When you try two models on one shot, give both the same references and the same prompt, so the difference you see comes from the model.
Long continuous takes. The longer a single clip runs, the more chances the face has to change. Write a cut point into the shot rather than relying on one long take.
Side profiles. If a shot shows the character in profile, give the model a profile reference from the pack rather than front views only.
Two characters in one frame. Check these shots closely. If a face slips, generate singles and cut between them for the conversation feel.
Wardrobe and props. Do not rely on the references alone for clothes. If the character wears a specific color across a series, name it in every prompt.
Close-up references on Seedance. In our August 2026 tests they failed with E005 on every run, on 2.0 and 2.5. The credits come back automatically, but retrying the same image does not help: reframe it waist-up or wider and ask for the close-up in the prompt.
Compositing. Don't paste the reference face onto a clip that drifted. Regenerate from the pack instead.
FAQ
How do I keep the same character across multiple AI videos?
Generate one canonical still of the character, build a pack of three to five angles from it plus one waist-up shot, and use that pack as the reference input for every clip. On 8frame that's Seedream 5 Lite at 6 credits per image, then image-to-video on Kling 2.6 Pro at 47 credits (silent) or a reference-capable model such as Grok Imagine's reference mode (from 35) or Kling O1 Reference (73). Compare every clip with the approved still, and never re-roll text-to-video hoping for a match.
Which AI video model is best for consistent characters?
We have not measured which model keeps a face closest to its reference, so there is no ranking here. Choose by what the shot needs, then compare the takes with your approved still. If the shot can open on the still, use image-to-video: Kling 2.6 Pro is 47 credits for 5 silent seconds, 5 or 10 seconds long. If it can't, use a reference mode: Grok Imagine (1 to 10 references, 35 credits for 5 seconds at 480p or 48 at 720p, up to 10 seconds), Kling O1 Reference (up to 7, 73), Gemini Omni Flash (1 to 7, 85, up to 10 seconds), Happy Horse (1 to 9, 95 at 720p, up to 15 seconds), Seedance 2.5 (up to 9, 157 at 720p, up to 30 seconds, waist-up or full-length references only) or Veo 3.1 multi-reference (up to 3, 448 for 8 seconds at 1080p).
Does Wan 3.0 take reference images on 8frame?
Not on the canvas. Wan 3.0 animates one start image there, and Wan 3.0 Prime's reference mode takes a reference clip. Base Wan 3.0 reference-to-video (34 credits for 5 seconds at 480p, 68 at 720p, 135 at 1080p) is available only through the 8frame MCP connector, also from a clip.
Can I use a start image and reference images together?
Not on Seedance, Gemini Omni Flash or Happy Horse: with both attached, the references are used and the start image is dropped. Kling v3 takes only the start image.
Why does Seedance reject my character reference?
In our August 2026 tests, Seedance's filter rejected every reference image with a close-up face, on 2.0 and 2.5, with an E005 "flagged as sensitive" error. A waist-up shot of the same person, with the face about a quarter of the frame height, or a full-length shot passed. Retrying the same image does not help; ask for the close-up in the prompt instead. See Seedance flagged as sensitive (E005).
Is character consistency solved in 2026?
No model on 8frame guarantees it. Work still-first, keep shots short, cut often, compare every take with the approved still, and plan on regenerating a couple of clips per sequence. Side profiles, long continuous takes and shots with two characters are the ones to check first.
Sources
- Prices (Seedream 5 Lite, Nano Banana 2, ShotGrid, Kling 2.6 Pro, Kling O1, Grok Imagine reference mode, Gemini Omni Flash, Happy Horse, Seedance 2.5 and 2.0, Veo 3.1, Wan 3.0 reference-to-video): 8frame tool registry cost functions, read 2026-10-02
- Reference limits, start-image and end-frame rules, lengths, resolution settings, the start-image-or-references behaviour of Seedance, Gemini Omni Flash and Happy Horse, and Wan 3.0's inputs on the canvas: 8frame tool configuration and the settings 8frame sends to each provider, read 2026-10-02
- Output resolution of Kling 2.6 Pro (1080p) and Gemini Omni Flash (720p), which the node does not let you choose: 8frame output files, read-only, 2026-09-30
- Wan 3.0 reference-to-video through the MCP connector: 8frame MCP connector tool definitions, read 2026-10-01
- Seedance close-up face refusals (E005): 8frame tests on Seedance 2.0 and 2.5 at 480p, August 2026; details in Seedance flagged as sensitive (E005)
The whole chain runs on one canvas: generate the still, fan the angle pack out, send it to three video models side by side. Plans start at $19 a month, watermark-free on every tier, at 8frame.co/pricing.
Try it without signing up: the 8frame canvas opens on a real board. The canvas is free and unlimited; generation is paid from $19/month.