Sutra

AI ad generation: what it does well, where it breaks, and how to tell in three seconds

A working guide from a studio that ships AI-assisted ads every week. The four failure modes that make generated ads read as fake, the one production order that fixes most of them, and a scoring tool you can run on your own creative right now.

What is in here
  1. What is AI ad generation actually good at?
  2. Why do AI-generated ads look fake?
  3. The fix is an order of operations, not a better model
  4. Score a hook before you build the film
  5. So should you use AI to make ads at all?
  6. We have no threshold for how generic is too generic
The short answer

AI ad generation is very good at volume, variation and iteration, and reliably bad at physics, continuity and brand-specific taste. Generated ads read as fake for four repeatable reasons: objects perform events they could not physically perform, layers move against each other at false depth, the payoff leaks into the setup, and the type is unreadable at the size people actually watch. None of those are model problems. They are production-order problems, and they are fixed by deciding on paper before anything renders.

What you get out of this
  1. The four failure modes, each with the frame-level tell that gives it away
  2. A production order that puts every expensive decision before the expensive stage
  3. A scoring tool for a hook, which you can run on a cut you already have
  4. The honest limits: what we still cannot do, and what we get wrong

Where our own numbers come from

Measured in this studio
1000+finished creatives shipped across every brand we have worked onSutra Haus production log
72hfrom brief to a first finished cut, held as a standing commitmentSutra Haus
5independent checks a film clears before anyone sees it: message, visibility, physics, story, commoditySutra Haus
2xthe minimum surplus we generate before selecting - you pick a winner, you do not fix a loserSutra Haus
These four are ours: counted from this studio's own production log, not from a survey. Everything else on this page that carries a number links to whoever published it.

01What is AI ad generation actually good at?

It is good at the parts of the job that are arithmetic. Making forty versions of a line. Resizing a cut into nine formats without losing the crop that mattered. Filling a hole in a shot list on a Tuesday afternoon when the shoot was in March. Producing a plausible ten-second establishing shot of a place nobody can fly to this week.

That is not a small list. Before generation, the cost of a variant was a day. Now it is a coffee. When the cost of a variant collapses, the correct strategy changes: you stop defending your one good idea and start generating a surplus, then selecting from it. That single shift is worth more than any individual model.

Rework means picking, not re-directing.

The rule we run production on

Vote, then see

When an AI-assisted ad of yours underperformed, what did you change first?

The trap is assuming the same collapse happened everywhere. It did not. The cost of judgement did not move at all. Somebody still has to know which of the forty is the one, and why, and what it will cost you when it runs at scale against a real audience with real money behind it.

What moved, and what did not

The honest split
What moved, and what did not
DimensionGot cheapDid not get cheap
Making a variantMinutesNo. Knowing which variant
Filling a missing shotOften possibleNo. Knowing the shot is missing
Resizing and reformattingEffectively freeNo. Deciding what survives the crop
Copy linesUnlimitedNo. The one claim you can defend
Motion on a stillOne clickNo. Whether that motion is physically possible
A whole finished adLooks closeNo. Being right about the first three seconds
The left column is where a tool earns its price. The right column is where a person still does the work, and where the cost of being wrong is highest.

02Why do AI-generated ads look fake?

Four reasons, in the order they cost you money. We know them because we shipped all four, watched them die, and wrote each one down as a rule. The full frame-by-frame atlas of why AI ads look fake goes wider than these four. Every kill below is one of ours.

The four tells, in order of how often they kill a cut

Where it breaks
01Faked physicsa photographed object doingsomething it cannot doKILL02False depthtwo layers cut from the samepixels, moving apartKILL03Leaked payoffthe answer shown before thequestion landsKILL04Unreadable typelegible on your monitor,gone at 360 pixelsKILL
01Faked physicsa photographed object doing something it cannot dokill
02False depthtwo layers cut from the same pixels, moving apartkill
03Leaked payoffthe answer shown before the question landskill
04Unreadable typelegible on your monitor, gone at 360 pixelskill
Every one of these is visible in a single frame or a single second. None of them needs a focus group to detect.

Tell one: an object performing an event it cannot perform

We learned this the expensive way, on a shot of a photographed loaf pulling apart that cleared every automated gate the engine had and was killed on sight for behaving in a way bread does not. Nothing about the craft was wrong. The event was wrong, and the eye catches a wrong event faster than it catches a soft shadow, which is where the whole battery of physics tests we run on generated footage starts.

The rule that came out of it: a still may move as a card. It can push, settle, drift, or have light travel across it. It may not perform physics. Real physical events come from real footage or from true video generation, never from animating a photograph into an event it never had.

What a photograph is allowed to do once it starts moving

Legal, and not
What a photograph is allowed to do once it starts moving
DimensionAllowedWhat the eye reads
Push - the crop window walks across the frameYes. YesCamera movement. Nothing about the object changed.
Settle - a small scale-in that lands on the cutYes. YesA lens finding its mark.
Drift - sub-pixel linear travelYes. YesLife in a still. Almost invisible, which is the point.
Traveling light across the surfaceYes. YesThe room changed, not the object. The richest move we have.
A declared stylized tear or papercut splitPartly. If declaredDesigned animation. The viewer is told it is a drawing, and forgives it.
A realistic split, pour, tear or deformNo. NoAn event that did not happen. Caught in under a second, every time.
Parallax cut from the subject's own pixelsNo. NoA doubled object shearing against itself.
Bounce, spring or overshoot on any moveNo. NoPhysics applied to something with no mass. Reads as a template.
The line is not craft, it is honesty about the medium. A card can travel through space. A photograph of bread cannot become two photographs of bread.
The exception, added laterA still may split, tear or pour when the treatment is openly stylized and reads as designed animation - a papercut transition, for example. The kill is not the event. The kill is a photograph pretending the event was real.

Tell two: two layers cut out of the same pixels

The cheapest way to fake depth is to duplicate one plate into two layers and slide them against each other. We built it that way once. What the eye saw was not depth, it was a single object doubled and shearing against itself, and it died the moment it moved. A frozen frame of that shot looks completely defensible, which is why no frame-based check will ever catch it. The rule since: split along true depth boundaries only. The subject stays whole, the background is a separate plate with the space behind the subject filled in, and no two layers in relative motion ever come from the same pixels.

What false depth looks like from the side

Layer anatomy
BGBackground platea separate photograph, or filled behindSUBJSubject, kept wholenever split from itselfTYPEType, on its own planewith a real plate under itGRAINOne grain plate over everythingwhat makes three sources one image
The middle layer is where cuts get killed. If you duplicate one object's pixels into two layers and move them apart, the eye reads the doubling instantly - even when it cannot say what is wrong.

Tell three: the payoff arrives before the question

Every film worth watching poses one question a stranger wants answered, and quarantines the answer until the question has been made to matter. Generated cuts leak constantly, because a generator has no idea which of your frames is the answer. It will happily put the finished product in frame two, and the film is over before it started.

One refinement, learned the hard way in food work: the finished product as frame one, presented deliberately, is legitimate and often correct. What kills is the product as the second frame - stumbled past, mid-setup, neither hook nor payoff.

The same fifteen seconds, leaked and not leaked

Second by second
0s1s2s3s4s5s6s7s8s9s10s11s12s13s14s15sHookProduct, ea…Nothing left to wantHookQuestion sharpened, three timesProduct, as the answerLEAKED CUTHELD CUT
Read it as a list
AtChannelWhat happens
0.0sLeaked cutHook
1.4sLeaked cutProduct, early
3.2sLeaked cutNothing left to want
0.0sHeld cutHook
1.4sHeld cutQuestion sharpened, three times
9.2sHeld cutProduct, as the answer
Top row: the product lands at 1.4s, in the middle of the setup, and the question dies. Bottom row: the same asset, held to 9.2s, and everything before it is doing work.

Tell four: type that vanishes at the size people watch

You judge your ad at 1080 pixels wide on a color-managed monitor, sitting down. It runs at roughly 360 pixels on a phone in one hand, in daylight, at a bus stop. Two floors decide whether the words survive that: contrast measured under the worst pixel behind each letter, not the average, and time on screen of at least 1.2 seconds or 0.35 seconds per word, whichever is longer. Both are covered properly in the legibility floors for captions on a phone.

Try it: will your line survive the phone?

Type into it
Six coral Love Patches laid out in a grid on dark slate with dried root and green sprigs, used here as a background for testing caption legibility

Words-
Dwell floor-
Your hold-
Contrast-
The dwell floor is ours and it is deliberately blunt: max(1.2s, 0.35s per word). The contrast reading is a simplification of what we actually run, which measures the worst local background under every glyph, not a single average.

What comes out the other end

Ours, made this way
Ephoria - campaign ad
Hotel client - campaign reel
Beauty device - campaign ad
Fashion - studio demo
Every one of these ran for a paying brand or a studio demo. All of them mix real product photography with generated coverage. Tap any of them to see it full size with sound.

03The fix is an order of operations, not a better model

Every one of those four failures is cheap to prevent and expensive to discover. They are all decided before a single frame renders, and all found after, if you let them be. So the fix is structural: move every decision to the cheapest stage that can hold it. That is the whole argument for pairing a deterministic pipeline with AI generation.

Where that order came from

The order we run now was written after a production wave we scored at two out of a hundred, when every correction was landing on finished renders instead of on paper. The full account of that inversion is in why we get paper approved before pixels. What replaced it: understand what the assets can do, choose the story on paper, commit the sound before the first cut, then generate a surplus and pick from it.

The order we run, and where it costs nothing to change your mind

The pipeline
00Brand and adsstudywhat they alreadyrun, and how itsounds01Corpus censusevery usable asset,and what it can do ina film02Story menu, onpaper6 to 12 candidates,each with its hookand its soundCHECKPOINT A03Direction lockbeats with realdurations, soundtrackcommittedCHECKPOINT B04Mass variants,then selectgenerate a surplus,present the bestthree to fiveCHECKPOINT C
00Brand and ads studywhat they already run, and how it sounds
01Corpus censusevery usable asset, and what it can do in a film
02Story menu, on paper6 to 12 candidates, each with its hook and its soundCheckpoint A
03Direction lockbeats with real durations, soundtrack committedCheckpoint B
04Mass variants, then selectgenerate a surplus, present the best three to fiveCheckpoint C
The three marked stages are the only points where a human verdict is needed. A change at Checkpoint A costs minutes. The same change after a render costs a day.
The rule that makes it workNothing renders before Checkpoint A. Not a test, not a quick look, not a mood piece. Once frames exist, people defend them, and the argument stops being about whether the idea is right.

04Score a hook before you build the film

Most of a short ad's fate is decided in its first second, and you can test that second before committing to anything. Pick the family your hook belongs to, then check the gates honestly. If it does not clear, you have lost ten minutes instead of two days.

The hook scorecard

Score a hook
Family-
Gates-
Score-
The weights are ours and they are opinionated. Treat the number as a conversation, not a verdict. What matters more is which gate you could not honestly tick.

05So should you use AI to make ads at all?

Yes, in the way you use a lens. The question is never whether the tool touched the work. It is whether anybody decided anything. An ad made entirely by hand with no point of view fails exactly as hard as a generated one, and costs more. If you are weighing the routes on price, we broke down what AI ads, UGC creators and an agency each really cost.

Where your next ad should actually come from

Answer three questions
There is no branch here that ends in nothing but generation, and none that bans it either. Both of those are positions, not decisions.

Four things people say about AI ads that we have watched fail

Flip them
None of these are strawmen. All four were said to us by somebody paying for creative.

06We have no threshold for how generic is too generic

Run this on a cut you already have

Tick as you go - it remembers
0%
Five checks, roughly four minutes. Run them cold, on a phone, with the sound on. If you cannot honestly tick a box, that is the note - you do not need a second opinion to act on it.

Our own documented failures are not historical. The scale error in a hotel piece - a stone crest rendered at roughly eight feet when it is nowhere near that - went through three reviews before anyone caught it, because everyone was looking at the light. We have shipped animation so subtle that at phone size it read as a rendering mistake rather than a choice. And we have no proven threshold yet for how generic a piece has to feel before it stops working, only a running count and a stack of verdicts.

That last one matters. Anybody selling you certainty about AI creative is selling you something. What we can offer is a written rule for every failure we have had, and the discipline to run the check even when the frame looks good.

Questions people actually ask

Open what you need
Can AI make a whole ad by itself?

It can make something ad-shaped by itself. What it cannot do by itself is decide which of the forty things it made is worth your money, and that decision is most of the value. In practice the useful split is: the model produces material, a person decides the story, the order, the sound and the cut.

How do I know if an ad was AI-generated?

Look for the four tells in this piece: an object doing something physically impossible, two layers sliding apart at false depth, the payoff appearing before the setup has earned it, and type that dissolves at phone size.

The give-away is almost never texture. It is behavior. Watch what things do, not what they look like.

Is AI ad creative cheaper than a studio?

Per asset, dramatically. Per result, it depends entirely on whether anyone is selecting. Volume with no selection is the most expensive creative there is, because you pay for it twice: once to make it, and again in media spend while it fails quietly.

Will Meta or TikTok penalize AI-generated ads?

Platform policy in this area moves, so check the current policy pages before you rely on any answer, including this one. The practical risk has never really been policy. It is that the audience recognizes the tells before the platform does, and the cost shows up as a hook rate rather than a rejection.

What does a hybrid pipeline actually look like day to day?

A brand kit and a study of what the brand already runs, a written census of every usable asset, six to twelve story candidates on paper, one chosen, the soundtrack committed before any cut is made, then a surplus of variants and a selection. Generation shows up inside that as a way to fill named gaps, never as the starting point.

How many creatives do I need before I know something?

More than one and fewer than you fear. The number depends on your spend and your conversion rate, not on a rule of thumb, which is why we built a calculator for it rather than quoting a number here.

If you want to see the difference rather than read about it, the offer below is the whole argument in one object: send a link to your product, get back one finished ad, free, yours either way.

Where the numbers came from

  1. Ahrefs. Short vs long content in AI Overviews (174,000 pages) - used for the citation-length figure
  2. Nielsen Norman Group. How little do users read?

Every figure above links to the place it was published. Numbers marked as ours are measured inside this studio and we say so where they appear. We do not print a statistic we cannot point at.

Badal Kariwal

Runs Sutra Haus, a one-person ad studio that has shipped over a thousand finished creatives - film and stills - for DTC brands and hotels. Writes here about what the work actually taught him, including the parts that failed. The person who reads your brief is the person who builds the work. Send him something to make.

The version of this argument you can watch

Send a link. Get one finished ad back.

No call, no deck, no invoice. One finished cut built from your own product, inside three days, yours to run whether or not we ever work together. If the four tells in this piece show up in it, you will know we did not follow our own rules.

Replies within a day. Ad within three.
Read next