Sutra

How many creatives do you need? The arithmetic, and where it breaks

Every page that ranks for this question asserts a number and none of them shows a method. Here is the arithmetic, a calculator that runs it on your own spend, and an honest account of the four places the math stops working.

What is in here
  1. How many creatives should you test on Meta?
  2. A creative learns from conversions, not from impressions
  3. What happens when you split a small budget across twenty variants
  4. Work out your own number
  5. Where does the arithmetic stop working?
  6. So should you make more than you can measure?
  7. What we cannot tell you
The short answer

Take your monthly conversions and divide by about 64. That is how many creatives you can genuinely learn something from this month, at the size of difference most creative tests are really hunting: a 50% relative one. Chasing a 20% difference costs about 400 conversions per creative instead, which is worked out in why twenty creatives usually produce noise rather than a winner. At $10,000 in spend and a $40 CPA, 64 gives you four creatives, not twenty. Ship more than four if you can afford to. Just be clear that above that line you are shipping creative, not testing it.

What you get out of this
  1. The actual sample-size arithmetic, written out, so you can check it rather than trust it
  2. A calculator that runs it on your own spend, CPM, conversion rate and target CPA
  3. What happens to a $10,000 budget split across twenty variants, in one chart
  4. The four conditions under which this whole calculation stops being true
  5. Why we still build more than we can measure, and what that is actually for

Four figures this post rests on

The numbers under the argument
5%of Meta creatives clear the bar Motion calls a winner, and the rate barely moves between top accounts and average onesMotion
18ads a week produces roughly 0.9 winners a week; 4 a week produces roughly 0.2, at the same hit rateMotion
26.1%false positive rate when you stop a test the moment it looks significant, against the 5% you believe you are runningEvan Miller
2xthe surplus we build before selecting: ten statics made to send five, raised per campaign when the economics justify itSutra Haus production log
The first two are Motion's, from 578,750 Meta creatives across 6,015 advertiser accounts. The third is Evan Miller's. The fourth is ours, from our own production log. Every one of them links to its publisher.

01How many creatives should you test on Meta?

As many as your conversions will pay for, which for most accounts under $20,000 a month is between three and eight. The number is a function of your budget and your conversion rate, not of your ambition. A creative that has not bought enough conversions to separate itself from noise has told you nothing, however confident the dashboard looks.

Nobody publishes it this way. The pages ranking for this question assert a figure and move on: five per ad set, three to five a week, twenty a month. We went looking for an official platform number too. Meta's own Ads guide, checked on the day this was published, specifies formats, aspect ratios and file sizes, and states no recommended number of creatives at all.

What different budgets can actually buy you

Two different questions
What different budgets can actually buy you
DimensionCan learn fromWorth shipping anyway
$2,000 a month, $40 CPANo. Under 13 to 4
$10,000 a month, $40 CPAPartly. About 48 to 12
$50,000 a month, $40 CPAYes. About 1930 or more
$10,000 a month, $150 CPANo. About 13 to 4
$10,000 a month, $12 CPAYes. About 1320 or more
The left column is arithmetic: monthly conversions divided by 64. The right column is a judgement call, informed by Motion's volume finding rather than by any test of ours. We hold no outcome data of our own, so treat the right column as a position we argue, not a result we measured.

Four numbers the ranking pages assert, and what each one leaves out

Flip them
None of these is a lie. Each is a reasonable default with its conditions stripped off, which is what makes them useless the moment your account differs from the account the writer had in mind.

02A creative learns from conversions, not from impressions

This is where the rules of thumb go wrong. Impressions are cheap and they feel like data, so a variant with 40,000 impressions and two purchases reads as tested. It is not tested. The thing you are trying to measure is a difference in a rate, and a rate is only as certain as the number of events underneath it.

What your money is actually buying, in order

The chain
01Spendbuys impressions, atwhatever your CPM happens tobe that week02Impressionsbuy conversions, at whateverrate your offer converts03Conversionsbuy certainty, and nothingelse in the chain doesTHE BOTTLENECK04Certaintyis what you came to the testfor in the first place
01Spendbuys impressions, at whatever your CPM happens to be that week
02Impressionsbuy conversions, at whatever rate your offer converts
03Conversionsbuy certainty, and nothing else in the chain doesthe bottleneck
04Certaintyis what you came to the test for in the first place
Every link is a conversion of one currency into another, and only the third one is scarce. Most creative testing plans are written as if the first link were the constraint.

The standard sample-size formula for comparing two rates, at 95% confidence and 80% power, simplifies neatly when your conversion rate is small. The impressions cancel and you are left with a number of conversions per variant, which is the only form of the answer that is any use to a person planning a month of spend.

The whole calculation, in five lines

Copy it
conversions this month   = monthly spend / your CPA

conversions per variant  = 16 / (relative difference you want to catch)^2
                             a 50% difference  ->   64 conversions
                             a 20% difference  ->  400 conversions

creatives you can test   = conversions this month / conversions per variant
spend each one needs     = conversions per variant x your CPA

Worked, at $10,000 a month and a $40 CPA:
  250 conversions / 64 = 3.9 creatives
  and each of those needs about $2,560 before its number means anything.
The 16 is the standard two-sided approximation at 95% confidence and 80% power. It is an approximation and it errs optimistic: run the same inputs through Evan Miller's sample size calculator and you will usually be asked for more, not less.

Two honest caveats sit on that formula. It assumes each variant gets a fair share of the spend, which on an auction platform it does not. And it assumes you decide the sample size in advance, which almost nobody does. Both make the real number worse than the arithmetic says, never better.

03What happens when you split a small budget across twenty variants

Nothing dramatic. That is the problem. Twenty variants on $10,000 a month at a $40 CPA is twelve or thirteen conversions each, and at that volume the ranking you see is mostly the order in which luck arrived. You will still pick a winner off it, because the dashboard sorts the column for you.

Conversions each variant gets, against what a verdict costs

$10,000 a month, $40 CPA
2 creatives125644 creatives63648 creatives316420 creatives136440 creatives664
Conversions each one getsConversions needed to call a 50% difference
See the numbers as a table
Conversions per creativeConversions each one getsConversions needed to call a 50% difference
2 creatives12564
4 creatives6364
8 creatives3164
20 creatives1364
40 creatives664
The second bar in each pair is flat on purpose: the evidence a verdict requires does not get cheaper because you ran more variants. Only two of these five splits clear it.

The failure mode is not that you learn nothing. It is that you learn something false and then act on it with real money, which is the subject of why twenty creatives usually produce noise rather than a winner. Splitting further feels like more information and is less.

A creative that has not bought enough conversions to separate itself from noise has told you nothing, however confident the dashboard looks.

The rule this whole post is an argument for

Vote, then see

How many new creatives did you put live last month?

04Work out your own number

Put your four numbers in. The last output is a sanity check rather than a target: if the CPA implied by your CPM and conversion rate sits far above the CPA you are budgeting to, the creative count above is optimistic and you should plan against the implied figure.

How many creatives can you actually learn from?

Put your numbers in
Conversions a month at your target CPA-
Creatives you can actually learn from-
Spend one creative needs before its number means anything-
CPA implied by your CPM and conversion rate-
The 64 conversions per creative is the standard two-proportion sample-size approximation at 95% confidence and 80% power, sized to detect a 50% relative difference. Want to catch a 20% difference instead? Multiply the requirement by six and divide the creative count by six.

05Where does the arithmetic stop working?

In four places, and three of them make your real number smaller than the calculator says. Any page that gives you a creative count without naming these is selling confidence rather than method.

The four conditions that break the calculation

One per tab
The platform is not running your experiment for you

Delivery is not a randomized trial. The system concentrates budget on whatever looks promising early, so your variants do not receive equal spend and the ones that lose early may never accumulate enough events to be judged at all.

The practical consequence: your effective number of tested creatives is lower than the number you uploaded, sometimes much lower.

  • Plan against the spend a variant will actually receive, not the spend you divided by hand.
  • If you need a fair comparison, run a structured split test rather than reading the delivery report.
Only the fourth of these is an argument for more volume. The first three are arguments for less, and for spending the difference on making the few you run genuinely different from each other.

Four of ours. Count the swings, not the files.

Ours, and the point is uncomfortable
The first two are the same product, the same claim and the same world in two layouts. That is one swing rendered twice, and running them against each other would spend real money to learn which layout luck preferred. The third and fourth differ on subject, world, energy and point of view, which is what a second test actually costs to make.

06So should you make more than you can measure?

Yes, and the reason is not that measurement is overrated. It is that the hit rate does not improve with skill. Motion's read across 578,750 Meta creatives is that roughly 5% qualify as winners, and that top accounts and average accounts sit at about the same rate. What separates them is swings taken: 18 new ads a week surfaces roughly 0.9 winners a week, 4 a week surfaces roughly 0.2.

So volume is real leverage, and it is leverage on the production side rather than the measurement side. Which is why the shape of our own production is a surplus that gets culled, not a single idea that gets refined. The last creative act is selection.

Rework means picking, not re-directing.

The rule our production runs on

Four survivors from sets that were larger than this

What a surplus looks like
Ephoria - campaign ad
Beauty device client - campaign ad
Hotel client - campaign reel
SaaS - studio demo
Each of these came out of a batch built to be culled. The default multiplier is 2x - build ten to send five - and it is a configuration number we raise per campaign, not a doctrine. Tap any to watch it full size.

Two things follow. Ship above your measurement line if you can afford to, but stop calling the extras a test. And make the extras genuinely distinct, because volume made of one idea inherits that idea's fate. The production economics of doing that without the cost getting away from you are in what failed renders actually cost.

07What we cannot tell you

The one thing we cannot give you is the thing a creative volume plan most wants: evidence that our creative wins more often than somebody else's. No ad in our own corpus carries a return on ad spend, so we have not measured it and neither has anyone in this category who claims it. The full version of that admission sits with the citation audit in where hook rate and hold rate benchmarks actually came from.

The one number of ours worth planning aroundOf the films that clear every automated check in our pipeline, only 43% survive a human eye. Budget for a hit rate rather than for uniform quality. That is true of your testing plan too: the question is never whether some of them will fail, only whether failing is cheap.

Before you commit next month's creative plan

Tick as you go - it remembers
0%
Six checks. If you cannot tick the first three, the number of creatives is not your problem yet.

That limit is worth stating plainly because it changes what this post is for. The arithmetic above is not ours and is not in dispute: it is the standard sample-size calculation, and you can check it against any calculator you like. What we contribute is the production side, which is what it costs to put a genuinely different swing in front of the auction rather than another render of the same idea. Those are separate skills and only one of them is a claim about performance.

Questions people actually ask

Open what you need
How many creatives should I test on Meta?

Divide your expected monthly conversions by about 64. At $10,000 a month and a $40 CPA that is 250 conversions and roughly four creatives you can genuinely separate from each other. Below that you are ranking noise. You may ship more than four for coverage and freshness, and that is a defensible thing to do, but it is not a test.

How many creatives per ad set should I run?

Meta's own Ads guide states no number, which is worth knowing before you take one from a blog. The practical constraint is that every additional ad in a set divides the same conversion pool, and the auction will not divide it evenly. Two to four distinct concepts per ad set is a reasonable working shape for accounts under $20,000 a month.

How many new creatives should I test a month?

Testing and shipping are separate budgets. Test as many as your conversions can carry, which is usually three to eight. Ship as many as your production can carry, because winner rates are roughly flat across accounts and volume is the lever that moves winner counts. Just do not read a verdict off the ones you shipped for volume.

Is it better to test many small variations or a few big swings?

Big swings, until you have the conversion volume to resolve small differences. A 50% difference needs about 64 conversions per variant to detect. A 20% difference needs about 400. Most accounts cannot afford the second, so testing button colors of a creative concept spends money to produce a number nobody should act on.

Does creative diversity actually matter, or is it a platform talking point?

It matters for a reason that has nothing to do with the algorithm: your effective test count is the number of genuinely different ideas you ran, not the number of files you uploaded. We retired eight finished ads at once after realizing they were identical to a viewer on argument, subject, world, energy, point of view and borrowed genre. Eight uploads, one swing.

What if my conversion volume is too low to test anything?

Then move the decision earlier, where it is cheap. Judge creative on paper and at feed size before it runs, use a higher-frequency proxy such as hook rate to catch the truly dead openings, and accept that the in-market comparison will be qualitative for now. The numbers behind those proxies deserve their own scrutiny, which is in hook rate and hold rate.

For most accounts the honest number of creatives you can learn from in a month fits on one hand. The twenty in your plan were one decision, made once, rendered twenty times, and then ranked by a column that had no right to be sorted.

Where the numbers came from

  1. Motion. Creative Benchmarks 2026: winners are rare (578,750 creatives, 6,015 advertiser accounts, $1.29bn Meta spend) - published 17 April 2026; a winner is defined there as a creative spending at least 10x its account median and at least $500 in total
  2. Evan Miller. How Not To Run an A/B Test - the 26.1% false positive rate from stopping a test the moment it looks significant
  3. Evan Miller. Sample size calculator for A/B tests - used as a cross-check on the approximation in this post
  4. Meta. Facebook Ads guide: ad format specs and recommendations - checked on 1 September 2026; it specifies formats and file specs and states no recommended number of creatives

Every figure above links to the place it was published. Numbers marked as ours are measured inside this studio and we say so where they appear. We do not print a statistic we cannot point at.

Badal Kariwal

Runs Sutra Haus, a one-person ad studio that has shipped over a thousand finished creatives - film and stills - for DTC brands and hotels. Writes here about what the work actually taught him, including the parts that failed. The person who reads your brief is the person who builds the work. Send him something to make.

Start the count at one

Get one creative you did not pay for, and put it in the test.

Send a link to your product and we will build one finished ad from it, free, inside three days, yours to run either way. Give it the spend the calculator above says one creative needs, and you will have a real answer instead of a ranked list.

Replies within a day. Ad within three.
Read next