← Back to blog

How to Make a Street-Interview-Style Ad with AI

The 4-step workflow for AI street-interview ads: two-character consistency, ambient audio cues, question-hook structure, disclosure rules, and the cost math vs a two-person shoot.

The street-interview (or vox-pop) format is the classic "we asked 50 strangers one question" ad. Done with AI, you end up with a 20-to-30-second clip where an off-camera or on-camera interviewer puts a product question to a passerby, the passerby reacts, and the reaction sells the product better than any script could. This guide is the 4-step workflow: two-character consistency, the ambient audio that makes it feel shot on a sidewalk, the question-hook structure, and where the disclosure line has to sit. A two-person street shoot with a location, permits, and talent runs $800 to $2,500 for a day. The AI version costs $8 to $20 in model credits per finished variant.

TL;DR

Each variant costs $8 to $20 and takes under 30 minutes after the first pass. You can test six question-hooks for less than one hour of a location shoot.

Why the format still converts

The street interview works because it fakes objectivity. A scripted spokesperson is obviously paid; a stranger on a corner reacting to a product reads as an unfiltered verdict. That perceived neutrality is the whole mechanic, which is exactly why the disclosure rules below are not optional. You are producing testimonial-style creative, not a documentary, and the viewer has to be able to tell.

The hard part with AI has always been that vox-pop needs at least two recognizable faces holding steady across a back-and-forth. Single-avatar UGC only has to lock one identity. Here you are cutting between an interviewer and a subject repeatedly, and if either face drifts between cuts the illusion collapses. Identity locking is what made this format practical in 2026. For the broader single-avatar version, see the full AI UGC workflow.

The 4-step workflow

Step 1: Write the question-hook

No model yet. The interviewer's first line is the scroll-stopper. It has to be a real question a real person would stop to answer, and it has to pre-load the product angle without naming it.

Weak: "Have you heard of [brand]?" Nobody stops for that.

Strong: "What's the one thing you'd change about your morning coffee?" or "Be honest, how many skincare products are in your bathroom right now?" The question invites a confession, and the confession is the setup your product resolves.

For this guide we'll use a concrete example: a cold-brew concentrate brand.

Write three to six question-hooks before generating anything. The question is cheaper to test than the production, and it is the variable that moves performance most.

Step 2: Generate both characters

Model: Higgsfield Soul 2.0

Upload one front-facing reference portrait per character, one for the interviewer, one for the subject. Higgsfield locks each identity to its reference, so you can generate every shot in the conversation and keep both faces stable. Generate the interviewer and subject in separate sessions, each anchored to its own reference, then cut them together in Step 4.

Subject reaction prompt:

Man in his 30s, casual jacket, standing on a busy city sidewalk, looks
slightly off-camera as if answering an interviewer, says "Like... forty
bucks a week? That's actually bad." with a sheepish, self-aware laugh.
Vertical 9:16, overcast daylight, handheld feel, shallow depth of field,
background pedestrians softly blurred. Clean audio.

Interviewer prompt (if on camera):

Woman in her late 20s holding a small microphone toward camera, friendly
and curious expression, standing on the same city sidewalk, says "How much
do you actually spend on coffee every week?" Vertical 9:16, overcast
daylight, handheld feel, matching background blur. Clean audio.

Generate three to five variants of each line with slightly different expressions. The identity holds because the reference does not change. Keep the lighting words ("overcast daylight") identical across both characters so the two halves of the conversation look like the same place.

Step 3: Generate the street and ambient audio

Models: Seedance 2.0 or Veo 3.1 (ambient-audio shots), Kling 3.0 (fast cutaways)

A street interview lives or dies on sounding like a street. Silence reads as a studio. Use a native-audio model for at least the establishing shot and the product hand-off so you get baked-in ambience: traffic, footsteps, distant chatter.

Veo 3.1 for the establishing/ambient shot:

Wide handheld shot of a busy city sidewalk, mid-morning, pedestrians
walking past, light traffic sounds and street ambience. Vertical 9:16,
overcast natural light. Documentary, unpolished, 4 seconds.

For the product hand-off, route to Seedance 2.0 and upload the actual product reference so the concentrate bottle stays accurate through the motion:

[Product reference] cold-brew concentrate bottle passed from one hand to
another on a city sidewalk, close-up, natural daylight, vertical 9:16,
product stays in focus center frame, 3 seconds, handheld framing.

For quick cutaways (feet walking, a hand gesture, the subject glancing at the label) use Kling 3.0. At roughly 60 seconds per clip it is cheap enough to generate several and keep the two best. These do not need audio; the ambient bed from your Veo shot carries under them in the edit.

Step 4: Assemble, caption, disclose

Tools: 8frame Studio or any NLE

Cut it as a conversation. The rhythm that works for a 25-second vox-pop:

Seconds Shot Content
0 to 3 Interviewer, on camera The question-hook
3 to 8 Subject reaction The wince and the honest number
8 to 14 Cutaway + subject Product hand-off, subject reads label
14 to 20 Subject The math out loud, the turn
20 to 25 Product still + CTA "Link in bio" end card

Add captions in white with a thin black outline. Sound-off viewing is the majority on both platforms, and in a Q&A format the words are the whole payload.

Then place the disclosure. This is a staged format that presents an actor as a candid stranger, so treat it like testimonial-style creative: label it as AI-generated per platform requirements, and never let the framing imply these are real, unpaid public opinions. The AI-generated label goes on at upload; if your creative implies "real customer reactions," add an on-screen "dramatization" note as well. The full breakdown of when and how is in the AI ad disclosure guide.

Cost math

Two-person street shoot:

AI street-interview workflow:

The point is not that AI is cheaper per clip. It is that you can test six questions instead of committing a full production budget to one. See the AI UGC vs real creators playbook for when to route a winner into a real shoot.

Common pitfalls

Faces drifting between cuts. The single most common tell. Fix it by never changing the reference image for either character mid-project. Same portrait, every session.

Dead-silent street. A vox-pop with no ambience reads as green-screen. Generate at least the establishing shot with a native-audio model and lay that bed under the whole clip.

Question that names the product. "Have you tried [brand]?" is an ad; "How much do you spend on coffee?" is a hook. Keep the product out of the first line.

Skipping the disclosure. A format built on perceived neutrality is exactly the one regulators and platforms scrutinize. Label it.

FAQ

Can AI keep two different faces consistent in the same video?

Yes, as of 2026, using per-character identity locking. Generate each character in its own Higgsfield Soul 2.0 session anchored to a single reference portrait, then cut the two together in the edit. The model holds each identity to its own reference, so the interviewer and the subject each stay stable across every cut. The mistake to avoid is swapping reference images mid-project.

Do I have to disclose that a street interview is AI-generated?

Yes. Both TikTok and Meta require an AI-generated content label on synthetic media, applied at upload. Because the vox-pop format implies candid real-world opinions, you should also avoid any framing that presents the actors as genuine unpaid members of the public. If the creative reads as real testimonials, add an on-screen dramatization note. See our disclosure guide for specifics.

Which model gives the most realistic street ambience?

A native-audio model like Veo 3.1 bakes traffic, footsteps, and crowd murmur directly into the clip, which is the most convincing option for the establishing shot. Seedance 2.0 is the pick for the product hand-off because it keeps the real product accurate through motion. Use Kling 3.0 for silent cutaways and carry the Veo ambience under them in the edit.


Write your three question-hooks, load two reference portraits, and run Step 2 first, the characters are the hard part. Clone a UGC assembly template on 8frame's workflow library, or start building on the canvas where every model above lives in one place.

Related articles

use caseWhite-Labeling AI Video: How Agencies Deliveruse caseHow to Make an ASMR Product Video with AIuse caseHow to Make a Day-in-the-Life Video with AI

Make it
move.

Stay in the loop

Be the first to hear about our launch and get product updates