Sutra

Why AI ads look fake: a frame-by-frame failure atlas

Sixteen named tells, sorted by the part of you that catches each one: your body, your memory, your eye, your reading, your ear, your sense of story. Each with the symptom, the cause, and the fix.

What is in here
  1. Why do AI ads look fake even when the resolution is high?
  2. Six organs, sixteen tells
  3. Physics: one kill wrote the law
  4. Why do all AI ads look the same?
  5. Type and sound are the two tells you can measure
  6. How do I check a cut for all of this in five minutes?
  7. Every detector we built was confidently wrong
The short answer

Resolution is not the axis. AI ads read as fake because of behavior: objects perform events they could not perform, layers shear against each other at false depth, motion is real but invisible at feed size, type dissolves at 360 pixels, sound arrives late, and the film answers a question it never asked. Six different parts of a viewer catch these, and each catches a different class of error. Almost none of them are fixed by a better model.

What you get out of this
  1. Sixteen named tells, each with the frame-level symptom that gives it away
  2. The four kills that wrote our own rules, told as what actually happened
  3. Why a set of ads can pass every individual check and still read as one ad
  4. A scored diagnostic you can run on a cut you already have

Four numbers from our own production log

Ours, counted off delivered files
14/15of our own delivered static ads failed a legibility re-measure once we read the worst pixel under each letter instead of the averageSutra Haus production log
43%of the films that clear every automated gate we run are still killed by a human eye afterwardsSutra Haus production log
360pxthe width a feed tile actually renders at, which is where every tell in this atlas is decidedSutra Haus
5/100the score a client put on a finished twelve-second film of ours, judging it as a stranger to the brandSutra Haus production log
All four are ours. They were counted off files that had already been delivered, which is the only reason they are uncomfortable enough to be worth publishing.

01Why do AI ads look fake even when the resolution is high?

Because resolution is not what you are catching. You are catching behavior. A 4K frame of a bottle pouring wrong is more obviously wrong than a 720p one, because there is more evidence in it. Every tell below survives a resolution bump, and a few of them get worse.

This matters for what you do next. If you think the problem is fidelity, you upgrade the model and wait. If you know the problem is behavior, you change what you ask the model to do, or you change which shot carries the moment. One of those costs a subscription. The other costs an afternoon and works.

Five things people say about the AI look, and what we found instead

Flip them
None of these are strawmen. Every one was said to us by somebody paying for creative, and two of them were said by us before we knew better.

02Six organs, sixteen tells

Sorting the tells by what they look like produces a list nobody can use. Sorting them by which part of a viewer catches them produces a triage. Physics is caught by the body, before thought. Continuity is caught by memory, one shot later. Composition is caught by the eye. Type is caught by reading. Sound is caught by the ear, usually as a reason to leave. Story is caught by the mind, and always too late to save the ad.

The order the six organs fire in

The shape of it
01Physicsthe body, beforeany thoughtRE-ROLL02Compositionthe eye, insidethe firstfixation03Typereading, about asecond in04Soundthe ear, usuallyas a reason toleave05Continuitymemory, one shotlater06Storythe mind, andalways too late
01Physicsthe body, before any thoughtre-roll
02Compositionthe eye, inside the first fixation
03Typereading, about a second in
04Soundthe ear, usually as a reason to leave
05Continuitymemory, one shot later
06Storythe mind, and always too late
Roughly, and on a phone. The point of the order is economic: a failure caught early costs you the viewer before the later organs ever get a turn, so a beautiful third act cannot rescue a wrong first contact.

The atlas: sixteen tells, filed by the organ that catches them

Search the atlas
16 terms
The impossible eventPhysics
A photographed object performing something it never did: splitting, pouring, tearing. The compositor has to invent geometry the camera never recorded, and the eye reads the invention before it can say why.
Weightless contactPhysics
Two things touch and nothing about either changes. No compression in the cushion, no give in the cloth, no shadow tightening where they meet. Contact is a two-way event, and generated contact is usually one-way.
The free-floating shadowPhysics
A shadow that is present but detached: wrong direction, wrong softness, or offset by a fixed number of pixels rather than by a fraction of the object's own width. The commonest reason a composite reads as pasted on.
Spring and overshootPhysics
Any bounce or elastic settle on a move. A photographed card has no mass, so borrowed mass reads as a preset. We ease with smoothstep only, and camera ramps inside a cut are deliberately linear.
Sheared parallaxContinuity
Depth faked by duplicating one plate into two layers and sliding them against each other. What the eye sees is a single object doubled, shearing against itself. It survives a still frame and dies in motion.
Per-render driftContinuity
The same person, product or room changing slightly between shots because each shot was generated separately. A button moves, a label reflows, a hairline shifts. Nobody names it. Everybody feels the film is not about one thing.
The cut inside the shotContinuity
A source clip that is itself a compilation, with hard cuts hidden inside it. Take an in-point across one and your cut sheet says one shot while the film plays two. We now scene-detect every source before cutting.
Hidden animationComposition
Motion that is real, expensive and covered, by a type layer or by its own subtlety. At 360 pixels it does not read as restraint. It reads as a rendering bug, which is the worse of the two readings.
Everything moves, nothing changesComposition
Objects animating busily while the light and the edit stay put. The films we build from do the opposite: elements are laid out once and then lit, and what moves is the light patch, the crop window and the cut.
Photo plus typographyComposition
An image with words on it and nothing assembled. This fails a commercial test rather than a realism one: if the buyer thinks they could do it in two minutes with a phone, the craft is invisible whatever it cost.
The sub-pixel stemType
An ultralight face at a small size whose strokes are narrower than one pixel column, so no stroke fully covers a pixel. Six to seven percent of its ink reaches full opacity. A size bump does not fix it. Weight does.
The scrim rampType
Pale display type over a photograph with a soft gradient scrim behind it. The scrim is a ramp, so it is weakest exactly where the type sits. Measured against the worst local ground rather than the average, most of these fail.
The silent first frameSound
Nothing on the soundtrack at frame zero, or a bed fading in. Across the films we have measured, sound is present in the first fraction of a second every time. A ramp spends the scroll decision in silence.
Treble as textureSound
A bright bed everywhere and no sound on the one sharp event in the picture. A blade, a click, a seal breaking. One sound on the one sharp thing beats a treble-rich mix, and costs nothing.
The leaked payoffStory
The answer visible before the question has been made to matter. In a making story, the finished thing appearing in shot one or shot two ends the film before it starts, however good the shot is.
The shuffleable filmStory
Beats that can be reordered without producing a lie. If your beat list still makes sense shuffled, it is an assortment. Passing looks like sentences that become actively false when you move them.
Every entry here came out of something we shipped, killed, or re-measured. Filter by an organ name to see one class at a time.

Two things follow from sorting them this way. The organs fire in order, so a physics failure spends the viewer before a story failure ever gets its turn. And they cost different amounts to repair: the body-level ones need the shot generated again, while the eye-level and reading-level ones are usually a rebuild of the last twenty minutes.

03Physics: one kill wrote the law

We built a shot of a photographed loaf pulling apart into two halves. Properly masked, color matched, grounded with a shadow derived from the element's own alpha. It cleared every automated check the engine had, because nothing in the engine knew what bread does. It was killed on sight, and the reason had nothing to do with craft.

Masahiro Mori's essay on the uncanny valley, translated for IEEE Spectrum, makes the argument that has held up best here: a near-miss reads worse than an obvious miss. A drawn hand asks nothing of you. A photographic hand with one wrong knuckle asks you to accept a person and then breaks the contract. Generated footage lives in that valley by construction, which is why small errors cost so much more than they look like they should.

That is not really how two pieces of bread separate.

The entire verdict on a shot that passed every automated check we owned

The crumb-pull, in the order it happened

Step through the incident
Two halves of a loaf, cut out and pulled apart

Masked cleanly, color matched to the plate, grounded with a shadow derived from the element's own alpha rather than drawn by hand. On a still frame it was genuinely good.

Cost of a change here: minutes
The useful part is not the kill. It is where in the day the kill landed. Three hours in, on a finished render, is the most expensive place a taste decision can happen.
The physical-plausibility law, in one lineA still or a cutout may move as a card and may never perform a physical event it would have to invent geometry for. If the treatment is openly designed, the event is legal, because nothing is pretending. There is a whole post on that choice: when a declared style beats a bad photoreal.

04Why do all AI ads look the same?

Because sameness lives on axes nobody measures. Cut rhythm, transitions, type and music are easy to vary and easy to count, so they get varied. Argument, subject, world, energy, point of view and borrowed genre are hard to count, so they stay fixed, and those are the six a viewer actually sees.

We retired eight finished ads in one go over exactly this: they differed on every mechanical axis and were identical on all six perceptual ones. The whole account of that set, and what it takes to build ads that do not wear out together, is in what makes an ad resist fatigue.

What our old quality check measured, and what a stranger sees

Compare
What our old quality check measured, and what a stranger sees
DimensionOur old lint measured itA stranger sees it
Cut rhythmYes. YesNo. Not as difference
Transition paletteYes. YesNo. Not as difference
Type treatmentYes. YesNo. Not as difference
Music and tempoYes. YesNo. Not as difference
The argument the ad makesNo. NoYes. Instantly
Who or what is on screenNo. NoYes. Instantly
The world it happens inNo. NoYes. Instantly
Its energyNo. NoYes. Instantly
Point of viewNo. NoYes. Instantly
The genre it borrowsNo. NoYes. Instantly
We threw away the lint that measured the left column the same day we found this, and replaced it with one that scores the right column. No two pieces in a set may match on more than two of those six.

Sameness is the failure a maker is least able to see, because you built each one separately and remember the differences. A viewer meets the set at once and remembers nothing. There is more on how this happens shot to shot in per-render drift and brand consistency.

05Type and sound are the two tells you can measure

Most of this atlas is judgement. Two entries are not. Text legibility and first-sound timing are both numbers, and both are checkable on a delivered file in about a minute. We re-measured fifteen static ads we had already sent, rebuilt byte-identical, this time reading perceptual contrast against the worst local background under each glyph run instead of a mean, and again at 360 pixels.

A legibility re-measure of fifteen delivered statics

The numbers
141Delivered cards6018Text runs inside them
FailedPassed
See the numbers as a table
What we re-measuredFailedPassed
Delivered cards141
Text runs inside them6018
Ours, measured on files that had already shipped. Ten of the failing runs lost their stroke entirely at embed width, meaning the letters had physically dissolved. The contrast ratios we test against come from the W3C guidelines linked below.

Two patterns accounted for nearly all of it: pale display type over a photograph with a soft gradient scrim, and small letterspaced caps. The one card that passed was the only one that did not put type over a photograph. The failure was the house style, not one card, which is the only version of this finding worth having. More on the floors in legibility floors on a phone.

Four of our own statics, and what stops each being a photo with words on it

Ours, made this way
The test is commercial rather than aesthetic. If a buyer believes they could rebuild it in two minutes with a phone and a free design tool, the craft is invisible whatever it cost to make.

06How do I check a cut for all of this in five minutes?

Watch it cold, on a phone, with sound, at arm's length, once. Then watch it again and only look at the places two things touch. Contact is where physics is decided, and it is the fastest read in the frame. Everything else can wait for a second pass.

Four frames from our own work, and what we check in each

Scroll it
A gray Ephoria dream patch disc lying on a black reflective surface, its edge catching a soft highlightA glass dropper bottle with a white label standing on a dark slate tile inside a brass-edged tray on a pale plinthTwo wooden sun loungers with white cushions and dark red bolsters flanking a closed cream parasol, seen across a turquoise pool, with white terraced villas on a dry hillside behindA man in a black vest swinging a black kettlebell overhead outdoors in front of a white corrugated wall
A gray Ephoria dream patch disc lying on a black reflective surface, its edge catching a soft highlight
Frame 1
Where the object meets its own reflection

A reflective ground is the cheapest lie detector there is. The reflection has to agree with the object about position, angle and softness, and it has to fade at the rate the surface roughness implies. If the object drifts and the reflection does not follow at the same rate, you have two layers, not one scene.

A glass dropper bottle with a white label standing on a dark slate tile inside a brass-edged tray on a pale plinth
Frame 2
Where the object meets the surface

Look for two shadows of different softness: a tight, darker one right at the contact, and a wider ambient one further out. A single soft blob under the whole object is the pasted-on look. The offset should scale with the object's own width rather than sit at a fixed number of pixels, which is why the same shadow recipe looks wrong when you change the size of the thing.

Two wooden sun loungers with white cushions and dark red bolsters flanking a closed cream parasol, seen across a turquoise pool, with white terraced villas on a dry hillside behind
Frame 3
Where the shadows disagree with the sun

One sun, one shadow direction, one shadow length ratio across the whole frame. Generated environments break this quietly: a parasol that shades nothing, a lounger whose shadow points a few degrees off its neighbor's. Also check that the water distorts what is under it and reflects what is above it, in the same shot.

A man in a black vest swinging a black kettlebell overhead outdoors in front of a white corrugated wall
Frame 4
Where mass is being claimed

A heavy object bends the person holding it. Look at the wrist, the shoulder and the opposite hip, not at the object. If the body would look the same holding nothing, the weight is a costume. This one is a favorite of ours because it is the tell you can check on a still, and almost nobody does.

None of these frames is a failure. They are the kinds of frame where a failure would be visible, which is why they are worth stopping on.

The motion nobody could see

A set of our films carried image animation that was, in the delivered file, covered by black type plates. The motion was real, it cost money, and at feed size it was invisible. The verdict was that viewers would read it as a mistake rather than as restraint, and that is right. A viewer cannot tell subtle from broken.

What we animated, against what the viewer received

Second by second
0s1s2s3s4s5s6s7s8sPlate settles inLight travels across the packSlow drift upType plate over all of it. Nothing perceptible.BUILTDELIVERED AT 360PX
Read it as a list
AtChannelWhat happens
0.0sBuiltPlate settles in
2.4sBuiltLight travels across the pack
5.2sBuiltSlow drift up
0.0sDelivered at 360pxType plate over all of it. Nothing perceptible.
This now has its own check, run by somebody who sees only the delivered bundle and never the build. Motion that cannot be perceived at feed size does not count as motion.

Invisible animation is worse than none.

The rule we wrote after paying to render motion nobody could see

07Every detector we built was confidently wrong

Every automated substitute for judgment we have built has failed toward a confident wrong answer. Four detectors, five limb measurers, a cut counter wrong by nine times in one direction, and a frame-delta gate that ranked our accepted films below our rejected ones and was withdrawn. The numbers behind that last one sit with the hook rate and hold rate benchmarks.

Which organ is catching your ad?

Score yourself
Five questions about a cut you already have. The score is a conversation starter, not a verdict. What matters is which question you could not answer honestly.

One limit runs under the whole atlas: not one film in the reference set carries a hook rate or a return on spend, so every tell here is a judgment about what breaks rather than a measurement of what converts.

There is no public field guide for video tells the way Wikipedia's editors built one for AI writing. That absence is most of the reason this page exists, and it is also a reason to argue with it. Sixteen is what we have written down so far, not a closed set, and the two we are least sure about are both in the composition row.

Questions people actually ask

Open what you need
Why do my AI ads look fake?

Almost always because of behavior rather than fidelity. Something in the frame does what it could not do: an object performs an event, a shadow floats free of its object, two layers shear apart at false depth, or a product changes slightly between shots. Watch your ad twice, and on the second pass look only at the places two things touch. That is where the eye is already looking.

Why do AI generated videos look fake even when the resolution is high?

Resolution adds evidence, it does not add correctness. A high-resolution frame of a hand holding a bottle wrong shows you more of the wrongness.

The tells that survive a resolution bump are all relational: contact, shadow, occlusion order, scale, continuity between shots. None of them is a property of one pixel, so no pixel count fixes them.

Why do all AI ads look the same?

Because the axes that are easy to vary are not the axes a viewer perceives. Cut rhythm, transitions, type and music are countable, so they get varied. Argument, subject, world, energy, point of view and borrowed genre are not countable, so they stay fixed, and those are the six that a stranger reads instantly. Vary the argument and the film stops belonging to the set.

What makes AI UGC ads look realistic?

Real contact and an ordinary performance. Hands on real objects, expression that has a cause in the frame, and a camera position that a person could have held. The register is set by audio more than picture: speech with gaps reads as a person talking, scripted voiceover with no gaps reads as a brand film. There is a whole piece on the hardest version of this, hands, products and the contact problem.

Is AI slop a real thing, or just a complaint?

It is a real and specific thing, and it is not the same as low quality. The complaint people are making is about the absence of a decision: work that could have been made about anything, for anyone, and shows no evidence that a person chose one thing over another. That is why volume alone reads as slop even when each individual asset is clean.

Can I fix these tells in post, or do I have to re-generate?

Appearance is fixable in post. Behavior is not. Grade, grain, crop, speed and sound can all be repaired after the fact. Contact, deformation, occlusion order and depth are baked into the pixels and need the shot made again. We wrote the full test battery for this at what video models still break.

A stranger catches a fake ad by noticing that nothing in the frame had to be decided by anyone. Every tell in this atlas is a decision that was left to a machine, and the fix in every case is to make it yourself, on paper, before anything renders.

Where the numbers came from

  1. IEEE Spectrum. The Uncanny Valley (Masahiro Mori, translated by MacDorman and Kageki) - the original argument that a near-miss reads worse than an obvious miss
  2. W3C. Understanding Success Criterion 1.4.3: Contrast (Minimum) - the 4.5:1 and 3:1 ratios our own legibility gate is built on
  3. Wikipedia. Signs of AI writing (community guideline) - the public field guide to AI tells in prose. Nothing equivalent exists for video, which is why we wrote this one

Every figure above links to the place it was published. Numbers marked as ours are measured inside this studio and we say so where they appear. We do not print a statistic we cannot point at.

Badal Kariwal

Runs Sutra Haus, a one-person ad studio that has shipped over a thousand finished creatives - film and stills - for DTC brands and hotels. Writes here about what the work actually taught him, including the parts that failed. The person who reads your brief is the person who builds the work. Send him something to make.