Sutra

How to vet an AI ad studio: nine questions and the answers to listen for

Nine questions that separate a studio with a process from a studio with a showreel. What a good answer sounds like, what a bad one sounds like, and our own nine answers filled in honestly, including the rows where somebody else is the better hire.

What is in here
  1. What you are buying is the behavior on the bad day
  2. The nine questions, and what a good answer sounds like
  3. Which reassuring answers actually mean the opposite?
  4. Why is the checking question worth two of the others?
  5. Now ask us the same nine questions
  6. So what do you do with the answers?
The short answer

Ask nine questions and listen for whether the answer contains a name, a number or a refusal. Vagueness is not politeness, it is the absence of a process. The nine: who actually builds it, what happens when a shot cannot be sourced, what a revision round costs, who owns the files, what they refuse to do, who checks the cut before you see it, what happens when round one is wrong, how they stop two clients looking alike, and what they will not promise.

What you get out of this
  1. Nine questions, each with the specific thing a good answer contains
  2. The four reassuring answers that mean the opposite of what they sound like
  3. A scorecard you can run during the call, not after it
  4. Our own nine answers, filled in honestly, including the ones where we are not your best option

What a showreel can tell you, and what only a conversation can

Compare
What a showreel can tell you, and what only a conversation can
DimensionA showreel answers thisOnly a question answers this
Can they make something that looks goodYes. Yes. This is the one thing it provesWhether the good one was the first attempt or the fifth
Do they have a recognizable styleYes. Visible in ten secondsWhether every client gets that same style whether it suits them or not
Will the person you met build your adNo. NoA name, and where that person will be on day two
What they do when a shot cannot be sourcedNo. NoWhether they refuse, substitute, or quietly generate it
What they refuse to doNo. Never on a reelThree refusals, each with a story attached
What happens when round one is wrongNo. Reels only contain round threeRepair or regenerate, and who pays for the difference
Who saw it before you didNo. NoWhether the person who built it is the person who checked it
A reel is round three of a project that had rounds one and two. It is real evidence of one thing and no evidence at all of five others.

01What you are buying is the behavior on the bad day

Every project has one hour in it where something cannot be done. The location has no photograph of the west-facing balcony. The model's hands do the wrong thing in every take. The one product shot you have is 900 pixels wide. What a studio does in that hour is the entire difference between the two quotes on your desk, and none of it appears in a reel.

The context you are vetting inside

Four numbers to hold during the call
5%of new creatives become winners, at the best accounts and the average ones alikeMotion, 550,000+ ads
37point gap between how positive advertisers assume consumers feel about AI ads and how they actually feelIAB
43%of our own films that clear every automated gate survive the first human viewingSutra Haus verdict log
5independent checks a film clears here before any client sees itSutra Haus
Two of these are external and linked. Two are ours, counted off our own production and verdict logs rather than surveyed.

The volume evidence makes this sharper. Motion's benchmark set of more than 550,000 ads found that the best-performing accounts and the average ones win at roughly the same rate. The top account wins more often because it ships more genuinely different swings, not because its craft is better. What you are hiring is a rate of good decisions under pressure, repeated weekly. So ask about the pressure.

A studio with no refusals has no standard. Refusals are the cheapest evidence of a point of view that exists.

why question five is the one we would ask first

There is a second reason to ask rather than watch. A reel is assembled for you, and craft is far more visible to a buyer than it is to a customer. What your customer notices is whether the first second gave them a reason to stay, and no reel can show you the ones that failed that test.

02The nine questions, and what a good answer sounds like

Ask them in this order. The first four are about the machinery, the middle three are about failure, and the last two are about taste. You do not need to be a video person to hear the difference. In a good answer somebody names a thing. In a bad answer somebody names a category.

Ask these, in this order

Nine tabs, nine questions
Who will actually do the building, and will they read my brief?

The person on the call is often not the person in the timeline, and the handoff is where a brief gets flattened into a template. On a short turnaround there is no time to recover from a flattened brief.

  • Good: a name, a role, and what happens if that person is ill. If part of it is subcontracted they tell you which part before you ask.
  • Bad: "our team of specialists." Follow up with: who reads the brief, and who cuts the film, and are those the same person?
Listen for
A human being with a calendar
The bad answers are all real answers that real buyers have accepted. None of them is a lie. Every one of them is a category standing where a specific should be.

Say them out loud rather than emailing them, and let the silence after each one run. The pause is data. Somebody who has thought about the question answers inside three seconds, because the answer is a thing that happened to them. Somebody who has not will reach for a category, and you can hear the reaching.

03Which reassuring answers actually mean the opposite?

Four answers get accepted because they sound like generosity or rigor. Each one is a real sentence somebody said to a buyer we later spoke to. None of them is dishonest. All four describe an absence.

Four answers that sound good and are not

Flip them
If any of these comes back, the follow-up question is the same every time: who, specifically, and what happens when they are wrong.

None of the four is disqualifying on its own. Good people say all four out of habit, because those are the sentences this industry taught them to say. What matters is the follow-up: who specifically, and what happens when they are wrong. Either a second, more specific answer exists underneath the first one or it does not, and you find out immediately.

04Why is the checking question worth two of the others?

No film leaves here until five separate verdicts come back: message, visibility, physics, story and commodity. Each one is run by a fresh reviewer who has seen only the delivered bundle, never the brief, never the build notes, never a previous verdict. A re-render voids all five, because a verdict belongs to the bytes that will actually ship.

Maker never checks maker, and the reason is a measurement rather than a principle. On the same films the build side's own scores ran between 36 and 76 points above the founder's. That gap is not dishonesty. It is what happens when the person who knows what a shot was meant to do watches the shot. There is a longer version of this in our deterministic pipeline write-up.

The five verdicts, and what each one is looking for

The shape of it
01Messagedo the words and thepicture make onehonest claimKILL02Visibilityis every motionperceptible at feedsize, or invisibleKILL03Physicsdesigned, or aphotograph pretendingan event was realKILL04Storydoes the questionhold, and does anearly frame leak itKILL05Commoditycould anybody elsehave made this exactpieceKILL
01Messagedo the words and the picture make one honest claimkill
02Visibilityis every motion perceptible at feed size, or invisiblekill
03Physicsdesigned, or a photograph pretending an event was realkill
04Storydoes the question hold, and does an early frame leak itkill
05Commoditycould anybody else have made this exact piecekill
Five separate reviewers, each given the delivered bundle and nothing else. Ask any studio to name their equivalent. The names matter less than whether the list exists and whether the maker is on it.

Score the studio you are talking to

Score the studio
Five of the nine, scored while the call is still fresh. The score matters less than which question you could not answer, because that is the one you did not really ask.

What the checks cannot do is tell you the ad is good. Fewer than half of our films that pass every automated gate survive the first human viewing, and that is what the gates are for. They clear the floors so the argument can be about whether the thing works, which is the only argument worth paying a studio to have. The money side of that is in what failed renders actually cost.

05Now ask us the same nine questions

A vetting post that cannot be pointed at its author is an advertisement. So here are our nine answers with the cost of each one written next to it. Three of these rows are reasons to hire somebody else, and those are the rows I would read first.

Sutra Haus, answering its own nine questions

Filled in honestly
Sutra Haus, answering its own nine questions
DimensionOur answerWhat that answer costs you
Who builds itYes. One person. The person who reads your brief cuts your filmOne person is one point of failure. Six ads a week in five languages is a job for a shop with a bench
When a shot cannot be sourcedYes. We refuse and name the frame that would unlock itSometimes that frame is a photograph only you can take, which makes our problem your ten-minute job
RevisionsPartly. A defect is ours and does not count. A change of mind is a roundWe do not sell unlimited rounds. If five people sign off on your creative, that will feel tight
FilesYes. Delivered masters and the sources they were cut fromThe build engine behind them stays ours, so another studio cannot pick it up mid-campaign
What we refuseYes. Faked physical events, borrowed brand devices, two clients in one visual languageYou will occasionally be told no about a shot you asked for, and the legal alternative may be quieter
Who checks itYes. Five checker passes, fresh eyes, bundle only. Maker never checks makerIt adds a step before you see anything, so the first cut lands later than it technically could
When round one is wrongYes. Regenerated from a changed instruction, never patched, and never without your yesRegeneration costs real money and time, so being wrong is slower here than at a shop that paints over it
Two brands looking alikeYes. One visual language per brand, checked across six perceptual axesWe will not reuse the treatment you admired on another brand's ad, even when you ask us to
What we will not promiseNo. No hook rate, no cost per result, no promise that this idea winsIf you need a performance number in the contract, somebody will give you one. We think it is invented
Filled in honestly. We do not win every row, and the second column is not modesty, it is the actual cost of the position in column one.

One more admission, because it belongs in a vetting post rather than in a footnote: not one film in our own reference library carries a click-through rate or a return on spend, so every ranking in it is a judgment about craft rather than a measurement of what converts. Put that question to us the same way you would put it to anybody else on your shortlist, and compare the answers.

The honest external evidence points the same way. CreativeX measured a ten point rise in creative quality score against roughly a two percent fall in media cost across about 822,000 observations. Real, statistically solid, and small. Craft is a floor most spenders fail, not the thing you win with. Anyone selling you craft as the performance lever is selling past their own evidence.

Four films from three different brands

Ours, so you can check
Ephoria - campaign ad
Hotel client - campaign reel
Beauty device - campaign ad
Cafe - studio demo
Made by the same person in the same year. Look at whether they feel related, because that is question eight being answered without words. If they do, we broke our own rule and you should say so.

06So what do you do with the answers?

Do not score the call and pick a winner. Take the three vaguest answers and write them into the first scope as plain sentences. Who builds it. What counts as a round. Who checks it before it reaches you. A studio that meant its answer signs that paragraph without editing it. A studio that did not starts negotiating the wording, which is the whole test performed in one move and for free.

Four sentences worth more than a two-hour call

Paste this into the scope
WHO BUILDS IT
  <name> reads the brief and cuts the film. If any part of the work
  is subcontracted, we are told which part, before it starts.

WHAT COUNTS AS A ROUND
  A change of mind is a round. A defect is not a round and is fixed
  at no cost, because it was not our instruction that produced it.

WHO CHECKS IT
  Before we see a cut it is checked by somebody who did not build it,
  against checks named here rather than described as a process.

A SHOT THAT CANNOT BE SOURCED
  We are told, and told what would unlock it, before any substitute
  is generated, sourced or shot.
Written as the buyer, in the buyer's voice, so it can be pasted into a statement of work without translation. The wording matters less than watching what gets edited.

The nine, on a card you can take into the call

Tick as you go - it remembers
0%
Roughly twenty minutes if you ask all nine and let the silences run. The three you skip will be the three that matter in week four.
The cheapest vet of allBuy one small thing first and watch what happens when it goes wrong. The comparison of subscriptions, freelancers and studios covers what a first project should cost in each model, and what a brief has to contain covers your half of it.

Questions people actually ask

Open what you need
What questions should I ask an ad agency before hiring them?

Nine, in this order: who actually builds it, what happens when a shot cannot be sourced, what a revision round costs and what counts as one, who owns which files, what they refuse to do, who checks the cut before you see it, what they do when round one is wrong, how they keep two clients from looking alike, and what they will not promise. Listen for names and numbers rather than categories.

How do I know if an AI ad studio is any good?

Look for evidence of failure that they wrote down. A studio that can tell you about the ad it killed, and the rule that came out of it, has a process. A studio that only shows finished work has a portfolio.

Then check one structural thing: whether anybody other than the maker looks at the file before you do. That single answer separates a production line from a person exporting from a timeline.

Is a showreel enough to judge a creative studio?

It proves one thing well, which is that they can make something that looks good. It cannot tell you whether that piece was the first attempt or the fifth, who built it, what they did when a shot was missing, or whether the next client gets the same treatment. Judge the reel for taste and the conversation for everything else.

What does unlimited revisions actually mean?

Usually that round one was not designed to be right. Unlimited rounds move the cost from the studio's invoice to your calendar, and calendars are the expensive resource in a media buy. The better answer names a number of rounds, defines what a round is, and makes defects free because they were the studio's own mistake.

Should I hire a freelancer, a subscription or a studio?

It depends on whether your bottleneck is volume, taste or coordination. Volume favors a subscription, taste favors a person whose work you can point at, coordination favors a shop with producers. We wrote the three side by side, with prices and what each actually delivers, in the buying cluster of this journal.

Who owns the files when I work with an ad studio?

Get it in writing before the first invoice, and ask about four things rather than one: the delivered masters, the source clips, the layered project file, and the music license and what it covers if you re-cut next quarter. Everybody agrees you own the final video. The argument is always about the other three.

The best vetting question is the one that costs the studio something to answer honestly, and the cheapest vetting method is a small paid project that you watch go wrong. Everything above is a way of getting to that faster. The rule sitting behind row seven, re-roll and never repair, is the single answer that most changes what a project costs you.

Where the numbers came from

  1. Motion. Creative Benchmarks, 550,000+ ads across 6,000+ advertisers and about $1.3bn of spend - the hit-rate finding: top accounts and average accounts win at about the same rate
  2. IAB. The AI Gap Widens - 505 consumers and 104 executives; the gap between how advertisers and consumers feel about AI ads
  3. CreativeX. Creative Quality Score against media cost, 822,000 observations at 99% confidence - used for the point that the craft floor is real, measurable and small

Every figure above links to the place it was published. Numbers marked as ours are measured inside this studio and we say so where they appear. We do not print a statistic we cannot point at.

Badal Kariwal

Runs Sutra Haus, a one-person ad studio that has shipped over a thousand finished creatives - film and stills - for DTC brands and hotels. Writes here about what the work actually taught him, including the parts that failed. The person who reads your brief is the person who builds the work. Send him something to make.

The vet that takes no meeting

Send a link. Get one finished ad back.

Free, yours to run whether or not we ever work together, built inside three days from your own product. Run the nine questions on it afterwards. If the answers do not match what is written on this page, that is the most useful thing you will learn about us.

Replies within a day. Ad within three.
Read next