Sutra

How to review AI creative when you are not a video person

You know something is wrong with the cut and you cannot name it. Here is the vocabulary: how to watch it cold, how to tell a defect from a taste note, and the five questions that turn a shrug into a note somebody can act on.

What is in here
  1. What should you look for when reviewing AI-generated ad footage?
  2. Watch it cold, at phone size, before anyone explains it to you
  3. A defect note and a taste note are different objects
  4. What do you say instead of "I don't like it"?
  5. "Make it pop" is a real feeling with a real cause
  6. When should you trust a note you cannot justify?
  7. The notes you are always allowed to give
The short answer

Watch it once, cold, on your phone, at the size a customer will see it. Then write three things down before you write anything else: what question the ad poses by second three, the second at which that question gets answered, and whether any earlier frame gives the answer away. Nearly every useful note a non-video person has ever handed us was one of those three, phrased badly.

What you get out of this
  1. The cold watch: how to see an ad the way a stranger sees it, in about ninety seconds
  2. How to tell a defect note from a taste note, because they go to different people and cost different money
  3. Five questions that turn "I don't like it" into something a maker can act on
  4. What "make it pop" usually means, and how to find the cause under the feeling
  5. The seven notes you are always allowed to give, even when you cannot justify them

Four numbers that decide how a review should go

Why your discomfort counts
360pxthe width a feed tile actually renders at, and the width we judge every cut atSutra Haus
43%of films that clear every automated check we run still get killed by a human eyeSutra Haus production log
37 ptsgap between what ad executives think consumers feel about AI ads and what consumers reportIAB
3sby which a stranger has to be able to say what question the ad is posingSutra Haus
The first, second and fourth are ours, counted off delivered files and our own verdict log. The third is IAB's, from a survey of 505 consumers and 104 ad executives.

01What should you look for when reviewing AI-generated ad footage?

Three things, in this order. Does the first second pose a question. Is the answer withheld until it has been earned. Does anything on screen do something it physically could not do. The grade, the music and the font are all downstream of those three, and none of them rescues a cut that fails one.

Your job on this call is to be the first stranger, and you are better at that job than the person who built the file, because you have not spent two days inside it. The maker cannot un-see the intention. You can only see what is actually there. That asymmetry is the whole value you bring to the call, and it expires the second somebody explains the ad to you.

A dark gym in backlit haze, a man mid-skip with a jump rope, other people skipping behind him, and a small letterspaced wordmark set across the top of the frame
On a laptop, a wordmark set that small reads fine. In a feed tile it is a smudge with a shape. Our floor is cap height at 5% of frame height, measured off the delivered file rather than the design.

02Watch it cold, at phone size, before anyone explains it to you

The most common review mistake has nothing to do with taste. It is watching the cut on a 27-inch monitor, full screen, with sound, after a paragraph telling you what the idea is. That is four advantages your customer will never have. Give them all back before you form an opinion.

The ninety-second cold watch

One screen at a time
Step 1
Put it on a phone and make it small

Not full screen. A feed tile, roughly a third of the screen width, which is about 360 pixels on most phones. Hold it at arm's length in ordinary daylight.

This one change catches more defects than any other single move we know, because it is the only viewing condition that matches the buyer's.

Step 2
Watch it once, with sound, and do not scrub

One pass, at speed, no pausing, no going back. The moment you scrub you stop being a stranger and start being an editor, and strangers are the scarce resource here.

Step 3
Write three answers before you speak

What question did it pose by second three. At what second was that question answered. Did anything before that second give the answer away.

Write them, do not think them. A note that survives being written down is usually a real note.

Step 4
Only now, watch it again and look for causes

Second pass is for diagnosis, and it is where timecodes come from. Your reaction came from the first pass. Do not let the second pass talk you out of it.

1 / 4

Ninety seconds, once. If your reviewer has already read the concept note, they cannot run this and you need somebody else in the room.

The reason the order matters is that a hook can only be spent once. We test our own openings by building two seconds of three different versions, showing them to somebody who has never seen the film, and asking exactly one question. Then that person is used up for that film, permanently.

Ask someone who has not seen the film, and ask once. The second viewing is worthless, because the hook is spent.

From our own hook-testing protocol

The cold watch, as a form you can paste into an email

Copy it
COLD WATCH - fill this in before you say anything

1. The question this poses by second 3:
   ...................................................

2. The second at which that question is answered:
   ...................................................

3. Any earlier frame that leaks the answer? (which one)
   ...................................................

4. What the words and the picture claim TOGETHER in shots 1 and 2:
   ...................................................

5. Anything doing something it physically could not do:
   ...................................................

If line 1 is blank, stop here. There is no film yet, and every
other note you write will be a note about the wrong thing.

03A defect note and a taste note are different objects

This is the distinction that saves the relationship. A defect is a fact about the file: the caption is unreadable, the second shot is out of order, the logo is the wrong blue. A taste note is a fact about you: it feels flat, it feels cheap, it is not us. Both are legitimate. They travel to different people, they cost different amounts, and mixing them into one message is why review calls go badly.

Two kinds of note, and what each one costs

Sort your note first
Two kinds of note, and what each one costs
DimensionA defect noteA taste note
What it sounds like"The line at 0:04 is gone before I finish reading it""It feels a bit flat"
What it is a fact aboutThe fileYou
Who can settle itYes. A measurement, in two minutesNo. Only the person paying
What it costs to act onMinutes. It is a repair.A re-roll, sometimes the whole direction.
How it should travelA timecode and one sentenceA kill, with the reason attached
Can you be wrong about itYes. Yes, and the measurement will say soNo. No. It can only be expensive.
Who it actually goes toWhoever built the fileWhoever chose the direction
The practical rule: send defects as a list of timecodes, send taste notes as a kill with a reason. Never bundle them, because a maker reading a mixed list fixes the cheap things and quietly ignores the expensive one.

The cost column is the one clients under-read. In generated production a bad take cannot be nudged into a good one, which is why we run a hard rule that a shot is re-rolled and never repaired. A defect is a fix. A taste note is a new generation. Saying which one you are giving is the single most useful sentence in a review.

04What do you say instead of "I don't like it"?

Answer one of five questions instead, and hand over the answer rather than the verdict. Each one converts a feeling into a location. None of them requires you to know what a J-cut is, and every one of them produces a note a maker can act on the same afternoon.

Five questions, and the note each one produces

One per tab
What is this ad asking me to wonder about?

If you cannot finish the sentence "I wanted to know whether...", the film has no question, and everything after second three is decoration on a thing nobody is waiting for.

  • Good answer: "I wanted to know what she was about to put on her face."
  • Bad answer: "It's about the product being natural." That is a topic, not a question.
  • The note it produces: "There is no question in the first three seconds."
Answer these in order. In our experience the first two catch about half of everything that is ever wrong with a cut, and both can be answered by somebody who has never opened an edit timeline.

05"Make it pop" is a real feeling with a real cause

Makers roll their eyes at it. They should not. It is a genuine perceptual report from somebody who lacks the words, and it usually has one of about five causes, most of which are measurable in under two minutes. The job is not to translate the phrase. It is to go looking for the cause with the person who said it.

Five things clients say, and what they turned out to mean

Flip them
All five of these were said to us, in these words, by somebody paying for the work. In four of the five cases the person was right and we were the ones who had to go and measure something.

Two of those five are pure measurement. Contrast has a published threshold: W3C's minimum-contrast criterion asks for 4.5:1 on normal text and 3:1 on large text, and we hold delivery to the first and the 360-pixel embed to the second. Time on screen has a floor too. The rest of that arithmetic is in the smallest caption that still works on a phone.

06When should you trust a note you cannot justify?

When it arrives on the first cold viewing, before anybody has explained anything, and it is about whether you wanted to keep watching. That reaction is the one thing in the room that is genuinely hard to fake and impossible to reconstruct. Justification can come later, from the maker, whose job it is.

There is evidence you are closer to the audience than the people making the work. IAB surveyed 505 US consumers and 104 ad executives and found the executives overestimating how positive consumers feel about AI-generated ads by 37 percentage points, a gap that had widened from 32. Inside our own pipeline the same shape shows up: builders scoring their own films ran between 36 and 76 points above the owner's score on the same files. Self-scoring measures compliance with the brief. It does not measure taste.

What kind of note do you actually have?

Answer two questions
There is no ending here that tells you to keep quiet. The point of sorting the note is that each of the four travels differently and costs differently, not that some of them are invalid.

07The notes you are always allowed to give

These seven need no film vocabulary, no timecode and no apology. Every one of them names something a stranger experiences rather than something a craftsperson did, which is exactly the kind of thing you are in the room to catch. If a studio treats any of them as a nuisance, that is information about the studio.

Seven notes that are always fair

Tick as you go - it remembers
0%
Run these after one cold watch on a phone. Four minutes. Anything you tick is a sentence you can send exactly as written, and none of them require you to propose a solution.

None of that is a substitute for a process. Reviewing well is downstream of briefing well, and if you are inheriting cuts you cannot evaluate, the fault is usually upstream in how the work was briefed and QA'd. The mechanical half of this, the part a machine can run before a human ever looks, is in our pre-launch QA checklist for generated footage, and the argument for pushing taste into code as far as it will go is in turning taste into checks a machine can run.

The words, so you can use them in the call

Search it
8 terms
Cold watchReview
One viewing, at feed size, with sound, before anyone explains the idea. It is the only condition under which your reaction matches a customer's, and it can be run exactly once per person per film.
Defect noteReview
A statement about the file that a measurement could prove or disprove. Unreadable type, a shot in the wrong order, the wrong blue. Costs minutes to act on and should travel as a timecode.
Taste noteReview
A statement about your reaction that no measurement can settle. It feels flat, it feels cheap, it is not us. Always legitimate, and always more expensive than a defect, because it buys a new take.
LeakStructure
The payoff appearing before the film has earned it. Usually the finished product, stumbled past in shot two. The commonest structural fault in generated cuts, because a generator cannot tell which frame is the answer.
Dwell floorType
The minimum time a line of type stays fully opaque. Ours is 1.2 seconds or 0.35 seconds per word, whichever is longer. Fades do not count, because text at partial opacity is not readable.
HookStructure
Whatever makes a stranger's eye stop before they have decided to look. Judged in the first half second, at feed size, muted. A hook is a mechanism, not an adjective about the opening shot.
Second hookStructure
The device that holds a drifting viewer near the midpoint. Any film over roughly ten seconds needs one, and in several films we have measured it carries more production value than the opening does.
Re-rollProduction
Generating a shot again from scratch instead of correcting the one you have. In generated production a bad take cannot be nudged into a good one, so a direction note buys a re-roll rather than a repair.
Eight terms, which between them cover most of what gets argued about in a review. None of them requires you to know anything about editing software.

Questions people actually ask

Open what you need
What should I look for when reviewing AI-generated ad footage?

Watch it cold on a phone at feed size, then check three things. Does the first second pose a question a stranger would want answered. Is the answer held back until it has been earned. Does anything on screen perform an event it could not physically perform. Those three catch most of what goes wrong. Grade, music and font come after.

How do I give feedback on a video ad without sounding vague?

Give a location, not a verdict. A second, a shot number, or a line of type. "I stopped caring at 0:06" is a complete, useful note and takes no craft vocabulary at all.

Then say which kind of note it is. A defect is a repair and costs minutes. A direction note is a new take and costs a generation. Saying which one you mean prevents the wrong thing being fixed.

What do I say instead of "I don't like it"?

Say what happened to you and when. "I did not know what it was about by second three." "I saw the ending too early." "I would not have stopped for it." Each of those is actionable and none of them is a design instruction. Avoid proposing fixes, because a client-proposed fix is usually a guess about a cause you have not diagnosed yet.

Is it okay to reject creative I cannot explain?

Yes, if you do it early and honestly. In our own pipeline only 43% of films that pass every automated check survive a human eye, so unexplainable rejection is a normal and budgeted event rather than a failure of the review. What is not okay is rejecting late, or rejecting by asking for small changes you know will not fix it.

How many rounds of feedback should a video ad take?

Fewer than you think, if the expensive decisions happened on paper. We run three checkpoints - A paper, B direction, C picks - plus the Mockup Gate between B and C, so by the time you see finished variants rework means choosing between them rather than correcting one at the most expensive stage. Why rounds pile up in the first place, and the brief fields that stop them, is in how to brief and QA ad creative.

Should I show the ad to other people before I send notes?

Show it to exactly one person who knows nothing about the project, once, at phone size, and ask them one question: what did you think it was going to be about. Do not ask whether they liked it. Their attention is spent after the first viewing, so a second opinion from the same person later is not a second opinion.

The best reviewer this studio has ever had cannot name a single thing he likes. He can only name what is wrong, instantly, every time, and he is right often enough that his kill reasons became the data the whole pipeline is tuned against. You do not need vocabulary to be useful here. You need to watch it once, cold, at the size your customer will see it, and write down what happened to you.

Where the numbers came from

  1. IAB. The AI Gap Widens (505 US Gen Z and Millennial consumers, 104 US ad executives) - the 37-point perception gap, published 15 January 2026
  2. W3C. Understanding Success Criterion 1.4.3: Contrast (Minimum) - the 4.5:1 and 3:1 ratios our legibility gate borrows

Every figure above links to the place it was published. Numbers marked as ours are measured inside this studio and we say so where they appear. We do not print a statistic we cannot point at.

Badal Kariwal

Runs Sutra Haus, a one-person ad studio that has shipped over a thousand finished creatives - film and stills - for DTC brands and hotels. Writes here about what the work actually taught him, including the parts that failed. The person who reads your brief is the person who builds the work. Send him something to make.

Practice on something that costs you nothing

Get an ad to review that you did not have to pay for.

Send a link to your product and we will build one finished cut from it, free, yours to run either way. Then run the cold watch on it. If it fails any of the seven notes above, you will have found that out on our time rather than on your media budget.

Replies within a day. Ad within three.
Read next