Sutra

Keeping one person the same person across six shots

Every render is a fresh sample, so the face you loved in shot one is a sibling of the face in shot four. Here is what drifts, what actually holds it together, and the shot count at which you should stop cutting and film one take.

What is in here
  1. What actually drifts when you regenerate a person?
  2. Why does it drift on every single render?
  3. Five methods that hold, and what each costs
  4. How long can a shot run before drift shows?
  5. When should you stop cutting and shoot one take?
  6. The kill that taught us continuity is a direction problem
  7. What we still cannot do
The short answer

You cannot make a generator remember a person, because every render is an independent sample with no memory of the last one. What holds a character together is production discipline, not prompting: lock one reference frame and drive every shot from it, film the real person once if you possibly can, repeat wardrobe and light direction in every shot line rather than only the first, keep individual shots short enough that nobody gets to study a face, and cut on movement so the eye is busy at the join. Past roughly eight shots, one continuous take is the safer bet.

What you get out of this
  1. The nine things that drift, ranked by how fast a viewer catches them
  2. Why identical prompts give you siblings rather than the same person
  3. Five methods that hold, and what each one costs you
  4. The kill that taught us continuity is a direction problem, not a model problem

Where a face gives you away between shots

Tap the numbers
Close crop of a woman's face at a bathroom mirror, water and cleanser on her cheek, one hand rising into the bottom of the frame
This frame is from our own work, and it is real footage. The points mark the places drift shows up first when a person is regenerated shot to shot instead of filmed.

01What actually drifts when you regenerate a person?

Nine things, and they do not drift at the same speed. Hair part, eyebrow shape, small marks and jewelry go first, inside one or two renders. Face geometry, skin texture and apparent age drift slowly and get noticed late. Room and light direction drift instantly and are the easiest of all to control, because they are the only ones you can fully write down.

The ranking matters more than the list. A viewer is not auditing your ad. They form one impression of a person in the first shot and then check that impression, cheaply and unconsciously, at every cut. Anything that changes the silhouette breaks it. Anything that changes a detail they were never going to look at does not.

The slow drifts are the dangerous ones

Fast drift gets caught. Somebody in the review says the earring is gone, and you re-roll. Slow drift ships. Jaw width and apparent age move a fraction per render, so by the fifth shot the person is two years older and nobody in the room can point at the frame where it happened. It arrives on the viewer's side as unease rather than as an error, and unease is the expensive kind, because there is nothing to fix a note against.

Nine attributes, and whether a cut can hide them

Ranked by how fast it shows
Nine attributes, and whether a cut can hide them
DimensionHow fast a viewer catches itCan a cut hide it?
Hair part and hairlineImmediatelyNo. No. It changes the head's outline
Eyebrow shapeImmediatelyNo. No. Brows carry identity
Moles, freckles, scarsOn the second lookPartly. Only if the shot is under a second
Jewelry and nailsOn the second lookYes. Yes, and better still, remove them
Wardrobe detailImmediatelyNo. No, but it is fully writable
Jaw and face widthSlowly, by shot fourPartly. Partly. Vary the shot size and it hides
Apparent ageSlowly, and it feels like uneaseNo. No. This one has to be locked
Room and backgroundImmediatelyYes. Yes, if the location genuinely changed
Light directionFelt rather than seenNo. No. It reads as a jump in time
The right column is the useful one. Where a cut hides it, you have a scheduling fix. Where a cut does not, you need the attribute nailed down before you generate anything.
A woman in a white robe pressing both palms to her cheeks in a tiled bathroom, shot square to camera
A frame from a beauty device client's ad. Face large, centered, held, both hands on it. This is the most expensive kind of shot to hold consistently across an edit, and usually the most valuable one in the film. That trade is the whole subject of this piece.

02Why does it drift on every single render?

Because each render is an independent draw. The model is not editing the last clip, or remembering it, or even aware there was one. It samples a fresh video from your prompt, your seed and your reference, and everything you did not pin down gets filled in from its own prior.

Six renders, one person, no memory

What carries over, and what does not
01Your referenceone frame, or onefilmed person, heldoutside the model02Shot linewardrobe, light, lensand age, restated infullRESTATE EVERY TIME03Samplean independent draw,filling every gapfrom its own prior04Nothing persiststhe next renderstarts from yourreference again, notfrom this clipTHE GAP05Selectionkeep the takes thatmatch, re-roll theones that do not
01Your referenceone frame, or one filmed person, held outside the model
02Shot linewardrobe, light, lens and age, restated in fullrestate every time
03Samplean independent draw, filling every gap from its own prior
04Nothing persiststhe next render starts from your reference again, not from this clipthe gap
05Selectionkeep the takes that match, re-roll the ones that do not
The dotted middle of this chain is where every continuity problem lives. Nothing travels from one render to the next except what you carry there yourself.

The practical consequence is the one people resist: a take that drifts cannot be nudged back. You re-roll it. Repairing a generated take means paying at the most expensive stage to half-fix a thing you could have redrawn, and it is why re-roll, never repair is a budget rule in this studio rather than a taste one.

Identical prompts do not give you the same person. They give you siblings.

What per-render sampling means in practice

03Five methods that hold, and what each costs

None of these is a trick. Each one trades something away: freedom, coverage, or the ability to change your mind later. Run them in this order, because the cheap ones make the expensive ones unnecessary more often than you would expect.

The order we run them in

One at a time
Method 1
Lock one reference frame and never move it

Pick a single still of the person, at the angle and light you want to be the truth, and drive every shot from that same frame. Not from the best frame of the previous clip, which is the intuitive move and the one that compounds drift: each generation inherits the last one's errors and adds its own.

Cost: your reference angle constrains your coverage. A frontal reference makes profiles hard, and you will feel that in the edit.

Method 2
Film the person once and reuse them

If a real human can stand in a real room for twenty minutes, this ends the problem entirely. You then use generation for the shots you could not film, the environments you cannot reach, and the volume you cannot shoot. Casting evidence supports it independently: VidMob and TikTok, across 1,678 ads and 7.3 billion impressions, found everyday people 1.7 times more likely to hook, and direct-to-camera worth about 50 percent more hooking power.

Cost: a person, a room, an afternoon. Cheaper than the third re-roll.

Method 3
Restate wardrobe and light in every shot line

Most briefs describe the person once, at the top, and then describe only the action per shot. The model does not carry the top of your document into shot five. Every line repeats the whole identity: age, hair, wardrobe, jewelry or the absence of it, and which side the key light falls on.

Cost: repetitive, ugly documents. Write them anyway. This is the highest-yield unglamorous fix on the list.

Method 4
Give drift no time to be seen

Across 48 films we measured, the median mean shot length is 0.83 seconds, and product films cluster between 0.13 and 0.35. A face on screen for six tenths of a second is an impression. The same face for four seconds is an audit.

Cost: a fast cut has to earn its speed. Pace hides flaws, and it also hides information, so it is a real trade rather than a free win.

Method 5
Cut on movement, never on stillness

Join two shots at the moment something is traveling: a turn of the head, a hand crossing frame, a step. The eye is tracking motion at the join and is not comparing faces. Cutting between two static portraits is the worst possible place to put a join, and it is where most people put theirs.

Cost: it dictates what you generate. You need takes with movement at the ends, which means planning the join before the render.

1 / 5

Method two is the one nobody wants to hear and the one that works best. Twenty minutes of phone footage of a real person outperforms every prompting technique in this list.
The method behind the methodsLock everything, change one variable. It is the same rule that governs our locked-frame hooks: one variable changing is a slideshow, two variables out of phase is a rhythm, and three or more is drift. When a shot comes back wrong, do not adjust three things at once and re-roll. You will never learn which one mattered.

04How long can a shot run before drift shows?

Long enough for the action in it, and no longer. We cap picture beats at about 2.5 seconds and cards at 3.0, because a five-word lockup needs 1.75 seconds of dwell and a ceiling and a floor were pulling against each other with the reader in the middle. Those are working numbers, not laws.

We used to run a hard rule that no generated shot could exceed two or three seconds. We retired it. A generated shot that genuinely holds, because it has an internal arc or its own scene cuts inside the take, runs at whatever length the direction declares. The judge is the shot itself, watched cold at 360 pixels. Length is the pace of the action, which is the argument in duration is the action's pace.

The numbers we run shot length against

Measured off delivered files
0.83smedian mean shot length across the 48 films we measuredSutra Haus corpus measurement
2.5sworking cap on a picture beat, raised to 3.0 for a card that has to be readSutra Haus build spec
8shots of one generated person before re-rolls usually cost more than filming would haveSutra Haus production log
360pxthe width every shot is judged at, because that is the real embed width of a feed tileSutra Haus build spec
All four are ours, measured off our own delivered files and production log. The first two are descriptions of what works in feed, not laws of nature.

Six shots, one person, fifteen seconds

Second by second
0s1s2s3s4s5s6s7s8s9s10s11s12s13s14s15sOpen, faceHandsProductRe-hook, faceUse, over shoulderClose, face and cardFirst impr…Checked ag…Confirmed…THE CUTFACE LEGIBLE
Read it as a list
AtChannelWhat happens
0.0sThe cutOpen, face
2.2sThe cutHands
3.6sThe cutProduct
5.2sThe cutRe-hook, face
7.8sThe cutUse, over shoulder
11.4sThe cutClose, face and card
0.0sFace legibleFirst impression is formed here
5.6sFace legibleChecked against the first
12.6sFace legibleConfirmed, or not
The face is properly readable for about 4.6 of these 15 seconds, in three separate windows. Everything else is hands, product, room and motion. That is the budget you are actually managing.

05When should you stop cutting and shoot one take?

Around eight shots, in our experience, and earlier if the person is the argument of the ad rather than its presenter. The math is unkind. Every additional shot is another independent draw and another join to survive, so risk compounds while the audience's tolerance does not.

There is a cheaper move than either extreme, and most people skip past it. Reduce how much of the film needs the face at all. Hands, product, room, the back of a head, a voice over an unrelated shot: all of these carry a person perfectly well, and none of them can drift into being somebody else. The related problem of hands, and why they break for a different reason, is its own piece on hands, products and the contact problem.

Cut it, lock it, or film it

Answer two questions
No branch here bans generation and none of them recommends it blindly. The variable is how much of your ad's credibility is resting on one face.

06The kill that taught us continuity is a direction problem

Four takes of an influencer-style shoot died together in one review. Every one carried continuous performed expression and head-shakes with nothing in the frame causing them. The person in them was consistent enough. That was not the complaint.

The verdict on those four takesThe model was explicitly exonerated. The fault was direction. Nobody had decided what the face should be doing in each second, so the generator supplied the only thing it had a prior for: constant, ambient performance.

What came out of it is a gate we now run before anything generates: a per-shot, second-by-second, plain-English timeline of expression and action, headed by the fraction of seconds that are performed against the same fraction measured off the reference. Neutral is a legitimate state and usually the correct one. Expression has to have a cause in the frame.

Before you generate a person in more than three shots

Tick as you go - it remembers
0%
Six checks, about ten minutes, all of them on paper. Every one of these is cheaper before the render than after it, which is the entire argument.

07What we still cannot do

We cannot hold one synthetic person across a campaign the way a signed actor holds a campaign. Nobody can yet, whatever the demos suggest. And nothing in our reference corpus attaches a hook rate to a face, so the shot ceiling below is craft judgment rather than a performance finding.

There is a second reason to keep the shot count honest. Motion's benchmark set, 550,000 ads and more than 6,000 advertisers, reports about half of all creatives switched off before day 28. You are building a stream of character films, and a method that only survives with a week of babysitting per ad does not survive contact with that cadence.

The five words we use for this on a call

Search the vocabulary
5 terms
Reference lockMethod
One still, chosen deliberately, that every shot in a film is generated from. Never the best frame of the previous output, because that compounds each render's errors into the next one.
Per-render driftFailure
The identity difference introduced by each independent sample. It is not a bug in a model and it does not get fixed by a better prompt. It is what sampling is, and it accumulates.
Identity lineBrief
The full description of a person, age through key-light side, pasted into every shot line rather than stated once at the top of a brief. Ugly to read, and the cheapest fix available.
Cutting on movementEdit
Placing a join at the moment something is traveling through frame, so the eye is tracking motion rather than comparing two faces. The opposite, cutting between stillness, is where drift gets caught.
Re-roll, never repairBudget
A generated take that comes back wrong is regenerated rather than corrected. Repair happens at the most expensive stage and usually breaks a second thing, and every re-roll is a decision with a cost attached to it.
These are our own working terms, not industry standards. They are here because a shared word for a failure is what lets two people fix it without arguing about what they saw.

Four of ours where a person carries the film

Ours, made this way
Ephoria - campaign ad
Fashion - studio demo
Fitness - studio demo
Hotel client - campaign reel
Every one of these mixes filmed people with generated coverage. In each case the person was fixed first and the film was built around them, not the other way round. Tap any of them to watch it with sound.

Questions people actually ask

Open what you need
How do I keep a character consistent across AI video shots?

Drive every shot from one locked reference frame rather than from the previous output, repeat the full identity description in every shot line instead of only the first, keep individual shots short, and place your cuts on movement. If a real person can be filmed for twenty minutes, do that instead and generate everything around them.

Can I use the same AI actor across multiple videos?

Across a handful, yes, with a locked reference and disciplined shot lines. Across a campaign, expect a family resemblance rather than a person, and design the creative so that nothing depends on the audience believing it is one individual.

The safest campaign-scale version keeps the face in short windows and lets hands, voice, product and room do the continuity work instead.

Why does the face change between shots even with the same prompt and seed?

Because each render is an independent sample rather than a continuation. A seed fixes the noise, not the identity, and any change in duration, motion, framing or wording moves the result. Everything you did not specify gets filled in from the model's own prior, which is why the same prompt returns siblings rather than one person.

How many shots can one generated character survive?

In our own work, up to about eight before the accumulated drift starts costing more in re-rolls than a shoot would have cost. Earlier than that if the face is large in frame, if shots run long, or if the person is the credibility of the ad rather than its presenter.

Does character drift actually hurt performance?

We do not know, and neither does anybody publishing on it, because there is no primary study measuring continuity against in-market results. What we can say is that it is caught by viewers instantly, that it reads as carelessness rather than as style, and that fixing it is cheap on paper and expensive after a render.

Is it better to fix a drifted take or regenerate it?

Regenerate. Repairing a generated take means correcting at the most expensive stage, and a face repaired frame by frame usually breaks its own lighting on the way. Our rule is that a bad take is re-rolled rather than repaired, and that every re-roll is a decision someone makes on purpose with a cost attached.

Continuity is a scheduling decision, made on paper, about how much of your ad you are willing to rest on a face a machine has to redraw six times.

Where the numbers came from

  1. VidMob and TikTok. The Science of the Hook - 1,678 ads, 7.3bn impressions; used for the casting and direct-to-camera figures
  2. Motion. Creative Benchmarks 2026: winners are rare - 550,000+ ads across 6,000+ advertisers; used for creative lifespan and cadence

Every figure above links to the place it was published. Numbers marked as ours are measured inside this studio and we say so where they appear. We do not print a statistic we cannot point at.

Badal Kariwal

Runs Sutra Haus, a one-person ad studio that has shipped over a thousand finished creatives - film and stills - for DTC brands and hotels. Writes here about what the work actually taught him, including the parts that failed. The person who reads your brief is the person who builds the work. Send him something to make.

The test that settles it

Send a link. Get one finished ad back.

One finished cut from your own product, inside three days, free and yours to run either way. Watch it twice: once for the story, once only at the person. If they are not the same person in every shot they appear in, we did not follow the method in this piece.

Replies within a day. Ad within three.
Read next