Sutra

A pre-launch QA checklist for generated ad footage

Twenty-one checks in the order a person should run them, each one doable in under a minute, on a phone, by somebody who is not a video person. Plus the reason each check exists, which is usually an ad of ours that died.

What is in here
  1. In what order should you QA a video ad?
  2. Physics, depth and the leak kill a cut outright
  3. How do you check type on a phone?
  4. What does an automated version of this look like?
  5. Who should run the checks?
  6. All green means the floors cleared, not that the ad is good
The short answer

Run the checks in the order of what kills a cut outright, then what makes it weak, then what makes it forgettable. Physics and depth first, because those are unfixable and instant. Then story and the leak. Then type, frame, sound. Then message logic and brand last, because those are arguments rather than defects. Every check below is doable in under a minute, on a phone, by somebody who has never opened an editing timeline.

What you get out of this
  1. The full checklist, twenty-one items, grouped and in running order, that saves your ticks
  2. Why each item exists, which in most cases is one of our own ads that died
  3. What the automated half of this looks like when it runs, and what it refuses to judge
  4. Who should run the checks, and why it must not be the person who built the ad

Six of these checks, on one frame

Tap the marks
A still ad: a white foil pouch of collagen patches with a gold circular label propped upright on a sunlit travertine ledge, a serif headline reading Your shelf is a to-do list, This isn't set into the blurred stone above it, and the brand lockup in the top left corner
One of ours. Six of the twenty-one checks below are visible in this single frame, which is the argument for running QA on stills pulled from the cut rather than on the cut playing at full speed.

01In what order should you QA a video ad?

Kills first, then weaknesses, then arguments. Physics and depth are unfixable once shot and visible in a second, so they go first. Story and the leak are next, because they decide whether anything else matters. Type, frame and sound follow. Message logic and brand come last, because they are judgments rather than defects.

The pre-launch checklist, twenty-one items in running order

Tick as you go - it remembers
0%
About twenty minutes cold, on a phone, with sound, on a cut you have not watched today. Your ticks are saved in this browser so you can come back to it. If you cannot honestly tick a box, that is the note, and you do not need a second opinion before you send it.

Nearly every line there exists because something of ours failed on it. That is the only reason to trust a checklist: not that it is complete, but that each line was paid for. If you are doing this on the buying side rather than the making side, reviewing creative when you are not a video person covers the same ground from the other chair.

Run it cold, and run it once. The first watch is the only honest one you get, because after that you know where everything is and your eye stops behaving like a stranger's. If you need a second opinion, hand the file to somebody who has not read the brief instead of watching it again yourself.

02Physics, depth and the leak kill a cut outright

These three are first because they are not fixable in a grade or a caption. A photographed object performing an event it could not perform is caught by the eye faster than a soft shadow is. Two layers cut from the same pixels shear against each other the moment they move. And a leaked payoff ends the film in beat two, whatever happens afterward.

A hand tips a glass carafe of dark coffee toward a stoneware cup balanced on a stack of books, lit hard from one side, with the shadow of a bare branch across the wall behind
A pour is a physical event, so it comes from footage. The check is not whether the shot is beautiful. It is whether anything in frame is doing something a photograph cannot do.

The bread kill is the one I keep going back to: two halves of a photographed loaf pulling apart, passing every automated check we had, killed on sight because the craft was fine and the event was invented. It is told in full in why AI ads look fake.

The depth version of the same mistake

We built a depth effect once by duplicating pixels from a single plate into two layers and sliding them against each other. In a still frame it looked defensible. In motion the eye saw one object doubled and shearing against itself, and it died on the first watch. This is the class of defect a frame-by-frame check cannot catch, which is why the item on the list asks you to watch rather than to scrub.

03How do you check type on a phone?

Two floors and one habit. The line has to stay fully opaque for at least 1.2 seconds or 0.35 seconds per word, and it has to clear a contrast ratio measured against the worst local background under each run of letters, at delivery width and again at about 360 pixels. The habit is doing it on the delivered file rather than in the editor.

Type your own line and see whether it survives a phone

Type into it
Six coral Love Patches laid out in a grid on dark slate with dried root and green sprigs, used here as a background for testing caption legibility

Words-
Dwell floor-
Your hold-
Contrast-
The dwell floor is ours and deliberately blunt. The contrast reading here is a simplification of what we actually run, which measures the worst local ground under every glyph run rather than a single average across the string.

A mean is what makes this check fail silently. One of our headlines measured a perfectly respectable average and 1.17 to 1 against the actual worst ground beneath it, because the word sat across a doorway edge and a lit white blouse in the same line. When we re-measured fifteen delivered statics this way, fourteen of them failed, and sixty of seventy-eight text runs failed with them.

A man in a black vest swings a kettlebell overhead in a concrete yard against a white corrugated fence, with a small letterspaced wordmark centered along the top edge of the frame
The wordmark along the top is the check. Small letterspaced capitals are the second most common way type disappears at feed size, after pale type over a photograph.

The failure was the house style rather than one card, which is the useful and unflattering version of that finding. The full treatment, including why weight beats size, is in the legibility floors for captions on a phone, and where the interface eats your frame is in safe zones as constraints.

04What does an automated version of this look like?

About half of the list above is measurable, so about half of it should never reach a human at all. Here is a real shape of a run: the floors get measured on the delivered bytes, the failures are named with the number that failed, and everything that needs judgment is routed out rather than guessed at.

The measurable half, running on a delivered file

Watch it run
A trimmed transcript with the project renamed. The two lines worth stealing are the mean note and the two SKIP lines. A check that cannot decide something has to say so, because a check that guesses is worse than no check.

Two properties make that run worth having. There is no force flag, so nobody can wave a failure through at five in the evening. And a re-render voids every verdict, because a verdict belongs to the bytes rather than to the project. Sorting the measurable half from the rest is the whole subject of turning taste into checks a machine can run.

05Who should run the checks?

Not the person who made the ad. A maker reads their own intention into a frame and a stranger reads only what is there. Nothing here is packaged until five separate verdicts come back from reviewers who have seen the delivered bundle and nothing else: no build notes, no direction lock, no previous verdicts.

The five passes, and the exact question each one answers

Switch between them
Do the words and the picture make one honest claim?

Watch the first two shots and write down the single combined claim they make together, from the customer's chair rather than the brief's.

  • Negative words over the product are a negative claim about the product
  • Fear points at the status quo, never at the product experience
  • No clinical or invented claims anywhere, at any size, in any corner of the frame
Fresh eyes, bundle only, maker never checks maker. You can run all five yourself on somebody else's ad in about twenty minutes. Running them on your own is the one configuration that does not work.

One more property is worth copying. A re-render voids all five verdicts, with no exceptions and no partial credit, because a verdict belongs to a specific file rather than to a project. It sounds pedantic until the first time somebody re-exports at the last minute to fix a caption and quietly ships a version nobody checked.

06All green means the floors cleared, not that the ad is good

This is the part a checklist post normally leaves out. Fewer than half the films that clear every automated check here survive the first human viewing, and that number is not a fault in the checks. It is what checks are for. They clear the floors so the remaining argument can be about whether the thing works.

Four numbers behind this list

Measured in this studio
5independent verdicts before a film is packaged, run by people who saw only the bundleSutra Haus
43%of films that clear every automated check survive the first human eyeSutra Haus verdict log
360pxthe width every legibility check is repeated at, because that is roughly a feed tileSutra Haus
14/15delivered statics that failed a contrast re-measure against the worst local groundSutra Haus legibility audit
All four ours, counted off our own delivered files and verdict log. The last one is a re-measure of statics that had already shipped, rebuilt byte for byte and then read against the worst local ground rather than an average.
The rule that stops QA making things worseA check is a floor, not a target. When a check fails, fix the thing rather than the number. The bed we tuned green and shipped anyway, and what the client said about it, is in sound is 40% of the ad.

The terms used in this checklist

Search the terms
6 terms
Dwell floorType
The minimum time a string stays fully opaque, computed from its own word count: at least 1.2 seconds, or 0.35 seconds per word, whichever is greater. Fades do not count, because text at partial opacity is not readable.
Legibility floorType
A two-part gate measured on the delivered file rather than inferred from a point size: cap height as a fraction of frame height, and perceptual contrast against the worst local background under each glyph run.
Full bleedFrame
Zero letterboxed or pillarboxed frames. Every frame fills the canvas edge to edge, and any fitted plate has to declare what fills the rest. Re-measured on the delivered file for flat edge bands.
Occlusion budgetLayout
A declared limit on how much of an element another element may cover, written when the composition is designed and measured after compositing. Ten percent hidden is a decision. Forty percent hidden is a bug.
The Stranger TestStory
Watching the full cut cold at feed size and answering in writing: what question does this pose by second three, when is it answered, does any earlier frame leak the answer, and what claim do the words and picture make together.
Message logicMessage
The rule that words and picture make one combined claim and must be judged together, from the customer's chair. Negative words over the product read as a negative claim about the product, however the copy was intended.
House vocabulary, offered because a shared name per idea is most of what makes a checklist run fast. The names matter less than everybody using the same one.

Questions people actually ask

Open what you need
What should I check before launching a video ad?

In order: physics and depth, then story and whether the payoff leaks, then type legibility, frame and bleed, then sound, then the combined claim the words and picture make, then whether the ad is distinguishable from everything it will run beside. The first three groups decide whether the ad is dead. The rest decide whether it is worth money.

How do I QA a video ad if I am not a video person?

Watch it on a phone, cold, at arm's length, with sound, without reading the brief first. Then answer three questions in writing: what is this asking me, when does it answer, and what would I say this ad is claiming.

Those three catch more real defects than any technical pass, and none of them require you to know what a keyframe is.

How long should creative QA take?

About twenty minutes per cut for a person, once the measurable checks have already run. If a human review is taking two hours, most of that time is being spent on things a machine should have refused before render: text outside its plate, an unreadable label, a letterboxed frame.

What is the most common defect in AI-generated ad footage?

Behavior, not texture. Objects performing events they physically could not, and layers moving apart at false depth. Texture is the thing people expect to catch and the easiest thing to fix. Watch what things do rather than what they look like.

Should the person who made the ad run the QA?

No, and this is measurable rather than a matter of principle. On the same films, our own build-side self-scores ran between thirty-six and seventy-six points above mine. Self-scoring measures compliance with the brief. It does not measure whether the ad works on a stranger.

If every check passes, is the ad good?

It is launchable, which is a different claim. Fewer than half our gate-green films survive a human viewing. Passing means no floor was breached, so the argument left at the end is the only one worth having: does anybody stop for this.

A checklist takes every argument about a defect off the table. What is left to disagree about is the work. Everything above is downstream of one decision anyway, which is approving the story on paper before anything renders.

Where the numbers came from

  1. W3C. Understanding Success Criterion 1.4.3: Contrast (Minimum) - the 4.5:1 ratio our contrast check measures against at delivery width
  2. BBC. Subtitle Guidelines - the closest thing to a published standard for how long text needs to stay on screen
  3. Meta Business Help Center. About text overlays and the Safe Zone for ads in Stories and Reels - where the interface sits over your frame
  4. VidMob and TikTok. The Science of the Hook - 1,678 ads, 7.3bn impressions, including the static-logo finding

Every figure above links to the place it was published. Numbers marked as ours are measured inside this studio and we say so where they appear. We do not print a statistic we cannot point at.

Badal Kariwal

Runs Sutra Haus, a one-person ad studio that has shipped over a thousand finished creatives - film and stills - for DTC brands and hotels. Writes here about what the work actually taught him, including the parts that failed. The person who reads your brief is the person who builds the work. Send him something to make.

Run the list on something of ours

Send a link. Get one finished ad back.

One finished cut built from your own product inside three days, free and yours to run whether or not we ever work together. Take the checklist on this page to it. If a box will not tick, tell us which one and we will tell you what we measured.

Replies within a day. Ad within three.
Read next