Keeping one person the same person across six shots
Every render is a fresh sample, so the face you loved in shot one is a sibling of the face in shot four. Here is what drifts, what actually holds it together, and the shot count at which you should stop cutting and film one take.
What is in here
- What actually drifts when you regenerate a person?
- Why does it drift on every single render?
- Five methods that hold, and what each costs
- How long can a shot run before drift shows?
- When should you stop cutting and shoot one take?
- The kill that taught us continuity is a direction problem
- What we still cannot do
You cannot make a generator remember a person, because every render is an independent sample with no memory of the last one. What holds a character together is production discipline, not prompting: lock one reference frame and drive every shot from it, film the real person once if you possibly can, repeat wardrobe and light direction in every shot line rather than only the first, keep individual shots short enough that nobody gets to study a face, and cut on movement so the eye is busy at the join. Past roughly eight shots, one continuous take is the safer bet.
- The nine things that drift, ranked by how fast a viewer catches them
- Why identical prompts give you siblings rather than the same person
- Five methods that hold, and what each one costs you
- The kill that taught us continuity is a direction problem, not a model problem
Where a face gives you away between shots
Tap the numbers
01What actually drifts when you regenerate a person?
Nine things, and they do not drift at the same speed. Hair part, eyebrow shape, small marks and jewelry go first, inside one or two renders. Face geometry, skin texture and apparent age drift slowly and get noticed late. Room and light direction drift instantly and are the easiest of all to control, because they are the only ones you can fully write down.
The ranking matters more than the list. A viewer is not auditing your ad. They form one impression of a person in the first shot and then check that impression, cheaply and unconsciously, at every cut. Anything that changes the silhouette breaks it. Anything that changes a detail they were never going to look at does not.
The slow drifts are the dangerous ones
Fast drift gets caught. Somebody in the review says the earring is gone, and you re-roll. Slow drift ships. Jaw width and apparent age move a fraction per render, so by the fifth shot the person is two years older and nobody in the room can point at the frame where it happened. It arrives on the viewer's side as unease rather than as an error, and unease is the expensive kind, because there is nothing to fix a note against.
Nine attributes, and whether a cut can hide them
Ranked by how fast it shows| Dimension | How fast a viewer catches it | Can a cut hide it? |
|---|---|---|
| Hair part and hairline | Immediately | No. No. It changes the head's outline |
| Eyebrow shape | Immediately | No. No. Brows carry identity |
| Moles, freckles, scars | On the second look | Partly. Only if the shot is under a second |
| Jewelry and nails | On the second look | Yes. Yes, and better still, remove them |
| Wardrobe detail | Immediately | No. No, but it is fully writable |
| Jaw and face width | Slowly, by shot four | Partly. Partly. Vary the shot size and it hides |
| Apparent age | Slowly, and it feels like unease | No. No. This one has to be locked |
| Room and background | Immediately | Yes. Yes, if the location genuinely changed |
| Light direction | Felt rather than seen | No. No. It reads as a jump in time |

02Why does it drift on every single render?
Because each render is an independent draw. The model is not editing the last clip, or remembering it, or even aware there was one. It samples a fresh video from your prompt, your seed and your reference, and everything you did not pin down gets filled in from its own prior.
Six renders, one person, no memory
What carries over, and what does notThe practical consequence is the one people resist: a take that drifts cannot be nudged back. You re-roll it. Repairing a generated take means paying at the most expensive stage to half-fix a thing you could have redrawn, and it is why re-roll, never repair is a budget rule in this studio rather than a taste one.
Identical prompts do not give you the same person. They give you siblings.
What per-render sampling means in practice
03Five methods that hold, and what each costs
None of these is a trick. Each one trades something away: freedom, coverage, or the ability to change your mind later. Run them in this order, because the cheap ones make the expensive ones unnecessary more often than you would expect.
The order we run them in
One at a time04How long can a shot run before drift shows?
Long enough for the action in it, and no longer. We cap picture beats at about 2.5 seconds and cards at 3.0, because a five-word lockup needs 1.75 seconds of dwell and a ceiling and a floor were pulling against each other with the reader in the middle. Those are working numbers, not laws.
We used to run a hard rule that no generated shot could exceed two or three seconds. We retired it. A generated shot that genuinely holds, because it has an internal arc or its own scene cuts inside the take, runs at whatever length the direction declares. The judge is the shot itself, watched cold at 360 pixels. Length is the pace of the action, which is the argument in duration is the action's pace.
The numbers we run shot length against
Measured off delivered filesSix shots, one person, fifteen seconds
Second by secondRead it as a list
| At | Channel | What happens |
|---|---|---|
| 0.0s | The cut | Open, face |
| 2.2s | The cut | Hands |
| 3.6s | The cut | Product |
| 5.2s | The cut | Re-hook, face |
| 7.8s | The cut | Use, over shoulder |
| 11.4s | The cut | Close, face and card |
| 0.0s | Face legible | First impression is formed here |
| 5.6s | Face legible | Checked against the first |
| 12.6s | Face legible | Confirmed, or not |
05When should you stop cutting and shoot one take?
Around eight shots, in our experience, and earlier if the person is the argument of the ad rather than its presenter. The math is unkind. Every additional shot is another independent draw and another join to survive, so risk compounds while the audience's tolerance does not.
There is a cheaper move than either extreme, and most people skip past it. Reduce how much of the film needs the face at all. Hands, product, room, the back of a head, a voice over an unrelated shot: all of these carry a person perfectly well, and none of them can drift into being somebody else. The related problem of hands, and why they break for a different reason, is its own piece on hands, products and the contact problem.
Cut it, lock it, or film it
Answer two questions06The kill that taught us continuity is a direction problem
Four takes of an influencer-style shoot died together in one review. Every one carried continuous performed expression and head-shakes with nothing in the frame causing them. The person in them was consistent enough. That was not the complaint.
What came out of it is a gate we now run before anything generates: a per-shot, second-by-second, plain-English timeline of expression and action, headed by the fraction of seconds that are performed against the same fraction measured off the reference. Neutral is a legitimate state and usually the correct one. Expression has to have a cause in the frame.
Before you generate a person in more than three shots
Tick as you go - it remembers07What we still cannot do
We cannot hold one synthetic person across a campaign the way a signed actor holds a campaign. Nobody can yet, whatever the demos suggest. And nothing in our reference corpus attaches a hook rate to a face, so the shot ceiling below is craft judgment rather than a performance finding.
There is a second reason to keep the shot count honest. Motion's benchmark set, 550,000 ads and more than 6,000 advertisers, reports about half of all creatives switched off before day 28. You are building a stream of character films, and a method that only survives with a week of babysitting per ad does not survive contact with that cadence.
The five words we use for this on a call
Search the vocabularyReference lockMethod
Per-render driftFailure
Identity lineBrief
Cutting on movementEdit
Re-roll, never repairBudget
Four of ours where a person carries the film
Ours, made this wayQuestions people actually ask
Open what you needHow do I keep a character consistent across AI video shots?
Drive every shot from one locked reference frame rather than from the previous output, repeat the full identity description in every shot line instead of only the first, keep individual shots short, and place your cuts on movement. If a real person can be filmed for twenty minutes, do that instead and generate everything around them.
Can I use the same AI actor across multiple videos?
Across a handful, yes, with a locked reference and disciplined shot lines. Across a campaign, expect a family resemblance rather than a person, and design the creative so that nothing depends on the audience believing it is one individual.
The safest campaign-scale version keeps the face in short windows and lets hands, voice, product and room do the continuity work instead.
Why does the face change between shots even with the same prompt and seed?
Because each render is an independent sample rather than a continuation. A seed fixes the noise, not the identity, and any change in duration, motion, framing or wording moves the result. Everything you did not specify gets filled in from the model's own prior, which is why the same prompt returns siblings rather than one person.
How many shots can one generated character survive?
In our own work, up to about eight before the accumulated drift starts costing more in re-rolls than a shoot would have cost. Earlier than that if the face is large in frame, if shots run long, or if the person is the credibility of the ad rather than its presenter.
Does character drift actually hurt performance?
We do not know, and neither does anybody publishing on it, because there is no primary study measuring continuity against in-market results. What we can say is that it is caught by viewers instantly, that it reads as carelessness rather than as style, and that fixing it is cheap on paper and expensive after a render.
Is it better to fix a drifted take or regenerate it?
Regenerate. Repairing a generated take means correcting at the most expensive stage, and a face repaired frame by frame usually breaks its own lighting on the way. Our rule is that a bad take is re-rolled rather than repaired, and that every re-roll is a decision someone makes on purpose with a cost attached.
Continuity is a scheduling decision, made on paper, about how much of your ad you are willing to rest on a face a machine has to redraw six times.
Where the numbers came from
- VidMob and TikTok. The Science of the Hook - 1,678 ads, 7.3bn impressions; used for the casting and direct-to-camera figures
- Motion. Creative Benchmarks 2026: winners are rare - 550,000+ ads across 6,000+ advertisers; used for creative lifespan and cadence
Every figure above links to the place it was published. Numbers marked as ours are measured inside this studio and we say so where they appear. We do not print a statistic we cannot point at.
Send a link. Get one finished ad back.
One finished cut from your own product, inside three days, free and yours to run either way. Watch it twice: once for the story, once only at the person. If they are not the same person in every shot they appear in, we did not follow the method in this piece.
Replies within a day. Ad within three.