AI Advertising

AI UGC avatar: the same face and the same voice, scene after scene

“Filmed by someone like you” content is the best-performing format on TikTok and Instagram. Generating it with AI is easy; generating it with a person who stays the same from one video to the next is much harder. Here is the three-lock method, demonstrated on a real series with its full prompts — and the rules to follow to stay honest.

The ToolAcces teamOctober 8, 2026~13 min read

Why UGC, and why it falls apart with AI

UGC (user-generated content) means those videos filmed on a phone, facing the camera, in a bathroom, a car, a kitchen. No studio, no perfect light, no ad voice. That is exactly what makes them work: they look like what people already watch between their friends’ videos, not like an ad break.

Producing this format with real creators is expensive and slow: casting, shipping the product, script back-and-forth, usage rights. AI promises to do it in an afternoon. And it delivers… for one video.

The trouble starts with the second. The face drifts: the jaw softens, the hair changes shade, the gold chain disappears. The voice drifts: the accent moves, the pace speeds up, the timbre changes. Yet a UGC series relies entirely on seeing the same person again. When they change, viewers don’t think “oh, it’s AI”; they just drop off, without knowing why.

A UGC series is a person you recognise. If they change, everything collapses.

The result we’re aiming for

Here is a series produced in the ToolAcces Studio: one character, six situations. The set, light, lens and framing change every time. The face, beard, hair, t-shirt and chain do not.

Reference image of the avatar, grey backgroundReference
Terrace
The avatar on a metro platformMetro
Office
The avatar on a Haussmann balconyBalcony
The avatar holding an iced coffeeProduct
Six images of the same avatar, generated separately in the Studio. Two of them were then animated.

This result doesn’t depend on a miracle model. It comes down to one rule: what defines the person is written once and never moves; everything else changes freely. The method fits in three locks.

Lock 1: the identity block

The identity block is a paragraph that describes the person, and nothing else: age, face shape, eyes, hair (colour, cut, style), skin and texture, jewellery, clothing. Not a word about the set, the light or the framing: those change in every scene, it never does. Here is our avatar’s:

The series’ identity block
A handsome man in his late twenties, editorial model features, strong jawline, dark brown wavy hair slightly tousled, light stubble, warm olive skin, brown eyes, athletic build, wearing a plain white fitted t-shirt and a thin gold chain necklace

What it contains, and why

  • Features visible from afar: strong jawline, slightly tousled wavy hair, light stubble. These are what make the person recognisable on a phone screen.
  • Distinctive marks: the thin gold chain. A recurring detail is a powerful anchor, for the model as for the viewer.
  • Simple, fixed clothing: a fitted white t-shirt. The busier the clothing, the more likely it is to vary.

Approve it before anything else

The block will be copied into every prompt of the series. A mistake here (a forgotten “brown eyes”, a badly described haircut) is paid for as many times as there are scenes. Reread it, test it on two or three images, and only then freeze it.

The order of blocks in the prompt, which decides the render

A UGC avatar prompt is always assembled in the same order. It isn’t fussiness: the first words weigh more than the last in the model’s interpretation.

  1. Beauty descriptors, first

    “Handsome”, “editorial model features”. They set the person’s register before anything else.

  2. Texture and realism

    Visible pores, natural stubble, individual eyelashes, slight skin tone variation. This is what separates real skin from a plastic render.

  3. Camera and lens

    28 mm, 35 mm, 50 mm depending on distance. The wide angle gives the natural distortion of a phone held at arm’s length.

  4. The light

    The location’s light, named precisely: platform fluorescents, afternoon sun under a red awning, a desk lamp at night.

  5. The composition

    The framing (“from chest up”) and what can be glimpsed behind, blurred.

  6. “Shot on iPhone 17 Pro Max front camera”

    The line that pulls the whole render toward a believable selfie rather than a studio photo.

  7. The negatives

    The usual ones (cartoon, CGI, plastic skin, watermark), plus “ordinary-looking person” and “beauty filter”.

Why beauty before realism

It is the most counter-intuitive rule. If texture terms come first (“visible pores”, “no retouching”), they pull the whole render toward documentary, and the model outputs an ordinary person photographed realistically. Placed after the beauty descriptors, they produce the opposite: a model photographed realistically. Which is exactly what a good UGC creator does.

Here is the full prompt for the metro scene. Spot the blocks:

Prompt · metro scene
USE THE REFERENCE IMAGE FOR FACE IDENTITY ONLY. DO NOT REPLICATE THE BACKGROUND, FRAMING, LIGHTING, OR CAMERA ANGLE FROM THE REFERENCE.

A handsome man in his late twenties, editorial model features, strong jawline, dark brown wavy hair slightly tousled, light stubble, warm olive skin, brown eyes, athletic build, wearing a plain white fitted t-shirt and a thin gold chain necklace, relaxed confident expression, standing on a Paris metro platform, mouth closed.

ultra-realistic cinematic photography with natural imperfections
visible skin pores, natural stubble texture, individual eyelashes, slight skin tone variation
shot on a 28mm wide lens, natural smartphone perspective distortion
cool fluorescent platform lighting from above, realistic hard shadows
medium close-up from chest up, white tiled metro walls, RATP signage and a train softly blurred behind him

Shot on iPhone 17 Pro Max front camera

NEGATIVE: No cartoon, no CGI, no 3D render, no plastic skin, no beauty filter, no ordinary-looking person, no amateur styling, no text, no logo, no watermark
The avatar on a Paris metro platform

Realism is a matter of millimetres

The realism of an AI face is read in details we look at without knowing it: a pore, an eyelash, a tone shift at the corner of the eye. The shots below, generated in the Studio in 4K, show the level of detail you can reach when the realism block is placed right.

Face
Eye
Eyebrow
Skin, lashes, brows: natural texture is what separates believable UGC from a “too perfect” image.

Lock 2: the reference image

The identity block describes; the reference image shows. Together they leave very little room for interpretation. The reference is a simple image of your character: front-facing, shoulders up, on a neutral background, with no distracting set.

The line that changes everything

Attaching a reference has a side effect: the model tends to copy the whole image, grey background and framing included. So every scene prompt starts with this line, in capitals:

Put at the top of every scene prompt
USE THE REFERENCE IMAGE FOR FACE IDENTITY ONLY. DO NOT REPLICATE THE BACKGROUND, FRAMING, LIGHTING, OR CAMERA ANGLE FROM THE REFERENCE.

Without a reference: find the face first

If you start from scratch, don’t look for the perfect face on the first try. Generate a broad batch, with no set constraint, from the identity block alone. Keep two or three, not one: a face that looks ideal as a still sometimes collapses once animated (frozen expression, a smile that distorts the jaw). Test the animation, then choose.

Never someone else’s face. Don’t use a real person’s photo without their written consent, and don’t try to look like a celebrity. An avatar generated from scratch is yours; a borrowed face exposes you.

Lock 3: the voice block

It is the piece nobody guesses, and the one that decides. A viewer forgives a slight change of light; they don’t forgive a voice that changes. The voice block is the audio twin of the identity block: written once, copied word for word into every talking video prompt.

Example voice block
VOICE: native English speaker, natural General American accent (not British, not Australian). Man in his late twenties, warm medium-low voice, slightly husky. Speaks at the pace of a voice note to a close friend: relaxed, small natural pauses, no rush. No vocal fry, no influencer cadence, no news-presenter voice, no radio announcer tone.

What it must name

  • The language and the precise region. “English” isn’t enough: the model may pick a British or Australian accent. “General American accent” closes the door.
  • Age and timbre. Low, medium, slightly husky.
  • The pace, compared to a real situation. “Like a voice note to a friend” means more to the model than “medium speed”.
  • What to avoid. No vocal fry, no influencer cadence, no presenter voice. That is what separates UGC from a radio ad.

The only thing that varies: the acoustics

From one scene to the next, you add a single sentence at the end of the block: the location’s. A voice with no echo in a bathroom sounds fake; so does a studio voice in the street.

Choosing sets without the series repeating itself

Once the person is locked, the whole point is to take them places. Sets are chosen by family rather than one by one: what you are really deciding is a register — intimate, outdoors, car. Golden rule: never two scenes from the same family in a row.

Intimate / mirror
Intimate / mirror

Bathroom, mirror selfie, morning routine. The closest register.

Home
Home

Kitchen, sofa, bedroom. The product in real life.

Outdoors, daytime
Outdoors, daytime

Terrace, street, rooftop. Air and natural light.

Car
Car

Passenger seat, parked. The quintessential UGC set.

Sport
Sport

Post-workout selfie. Energy and sincerity.

Going out
Going out

Café, restaurant, evening. The product in a social moment.

Vary the scale too: close-up, waist-up, full body. The set and framing change; the person, never.

Animating the avatar and making it talk

An image isn’t enough: UGC lives in video. There are two ways to bring the avatar to life, answering two different needs.

A silent shot: a single gesture

For shots without speech, the rule is that of any animation: one gesture, plus what doesn’t move. Here is the real prompt for the terrace scene.

Animation prompt · terrace
A handsome man sits at a Parisian cafe terrace, marble bistro table in front of him. Slow gentle push-in camera movement toward him. He picks up the espresso cup from the table with one hand, brings it to his lips, and takes a slow unhurried sip while looking slightly off to the side, then lowers the cup back down onto the saucer with a soft clink. Natural relaxed body language, no rush. Warm late afternoon sunlight, dappled shade from the red awning, realistic soft shadows. Ultra-realistic, natural skin texture, real phone camera look, mild grain. Silent, no dialogue, no added music.

A talking shot: the voice block goes in the video prompt

It is the most common mistake: writing the voice in the image prompt. An image has no sound; the voice block belongs in the video prompt, together with the line to speak, in quotes, in the target language.

Example · talking scene in a car
Handheld selfie video, phone at arm's length, the man from the reference image sitting in the passenger seat of a parked car, holding an iced coffee can. He looks into the lens and says: "Three things I wish I'd known before making iced coffee at home." Small natural head nods, one hand gesture.

VOICE: native English speaker, natural General American accent (not British, not Australian). Man in his late twenties, warm medium-low voice, slightly husky. Speaks at the pace of a voice note to a close friend: relaxed, small natural pauses, no rush. No vocal fry, no influencer cadence, no news-presenter voice, no radio announcer tone. Natural car interior acoustics, muffled street sounds outside.

No background music, no voice-over, no subtitles, no on-screen text.

With a HeyGen avatar, the voice lock takes another form: pick a voice from the library (or upload your own audio) and keep exactly the same one for the whole series. To go further on voices, see why an AI voice sounds robotic.

Building a 30- to 60-second video

A long UGC video isn’t a sixty-second shot: it is a sequence of segments, each with its own goal. Without that structure, you repeat the same hook-demo-call structure four times, and viewers leave halfway.

  1. Hook and context

    The most attention-grabbing moment, from the first second. The product can appear, but no sales pitch.

  2. Usage

    The product opened, applied, shown. Explain what it is and what it’s for.

  3. Proof

    A believable result, a demonstration, a common objection and its answer, a tip.

  4. Payoff and call

    The product as part of daily life, then a casual recommendation.

Each segment is generated separately, then edited back to back. Same character, same location, same light and the same voice block in all of them: that is what gives the impression of a single take filmed by the same person. For editing and subtitles, see how to edit and subtitle an AI video.

3locks: identity, reference, voice
7blocks per prompt, beauty first
4segments for a long video
60 sof continuous talking take with HeyGen

What kills a generated UGC video

  • Studio lighting. Nobody films their morning routine under three softboxes.
  • Poreless skin. The beauty filter is the signature of an image generated without care.
  • Impossible hands, a floating product, a warped label.
  • Background music or voice-over added by the model: always specify “no background music, no voice-over”.
  • On-screen text generated in the image: subtitles are added in the edit, never in the prompt.
  • Influencer speak. “Game changer”, “I’m obsessed”, “literally insane”: viewers recognise those tics in a second.
  • An identity block rewritten along the way. One more adjective, and it’s someone else.

The rules: an avatar is not a customer

This is the point you can’t get around. Classic UGC draws its strength from the fact that a real person really used the product. An AI avatar, by definition, hasn’t. Having it tell a lived experience (“I tested it for a month, my skin changed”) amounts to manufacturing a fake testimonial.

In Europe, unfair commercial practices law prohibits posing as a consumer and publishing fake reviews. In the United States, the FTC has banned fake testimonials since 2024, including AI-generated ones. The EU AI Act also requires, since August 2026, disclosure of certain synthetic content depicting people, and ad platforms have their own labelling rules.

What an avatar can do

  • Present the product, like a host would.
  • Demonstrate how it’s used, step by step.
  • Explain its ingredients, how it works, its limits.
  • Answer customers’ frequent questions.

What it must not do

  • Claim to be a customer, or to have used the product.
  • Announce a personal result (“I lost”, “my skin changed”).
  • Deliberately resemble a real person.

The obligations are detailed in what the AI Act changes for your creations. Good news: the “host” format works very well, precisely because it keeps the tone and settings of UGC without lying about who is speaking.

Automating: the avatar from your AI assistant

This whole method is available in the ToolAcces MCP connector, as a command: avatar-ugc. You describe the character in one sentence (“28-year-old man, dark hair, coffee lifestyle”); the assistant writes the identity block and submits it to you, then the voice block, picks sets without ever repeating a family, orders every prompt the right way, and announces the cost before launching.

If you have a reference image, it uploads it once and attaches it to every generation, with the capitalised line at the top of the prompt. Everything is explained in our MCP feature, and the same locking logic applied to products is detailed in the twelve-visual consistent campaign.

Frequently asked questions

Because every generation starts from zero and reinterprets your description. If the description of the person varies, even by one word, or is mixed with the set, the model rebuilds a slightly different face. The identity block, copied verbatim, and the reference image attached to every generation remove most of that drift.

With a voice block copied word for word into every video prompt: language and region, age, timbre, pace, things to avoid. Only the acoustics of the location change. If you use a talking avatar with a synthetic voice, pick one voice once and for all and never change it.

No. Passing off a generated character as a consumer who really used the product is a misleading commercial practice, and fake testimonials are explicitly targeted in both Europe and the United States. An AI avatar can present, demonstrate, explain; it must not invent a lived experience.

A video model generates shots of roughly 5 to 10 seconds. A 30- to 60-second video is built from segments edited back to back, each with its own goal. For a single continuous talking take, the Studio's HeyGen avatar goes up to 60 seconds.

No, and it is better to do without: first generate a batch of faces, keep one, and that image becomes the reference. Never use a real person's face without their written consent, nor a face deliberately resembling a celebrity.

Because UGC must look like what people film themselves: wide-angle lens, slight perspective distortion, arm's-length framing. That line pulls the render toward a believable selfie rather than a studio photo, which instantly gives the ad away.

Go further