How many creatives do you need? The arithmetic, and where it breaks
Every page that ranks for this question asserts a number and none of them shows a method. Here is the arithmetic, a calculator that runs it on your own spend, and an honest account of the four places the math stops working.
What is in here
Take your monthly conversions and divide by about 64. That is how many creatives you can genuinely learn something from this month, at the size of difference most creative tests are really hunting: a 50% relative one. Chasing a 20% difference costs about 400 conversions per creative instead, which is worked out in why twenty creatives usually produce noise rather than a winner. At $10,000 in spend and a $40 CPA, 64 gives you four creatives, not twenty. Ship more than four if you can afford to. Just be clear that above that line you are shipping creative, not testing it.
- The actual sample-size arithmetic, written out, so you can check it rather than trust it
- A calculator that runs it on your own spend, CPM, conversion rate and target CPA
- What happens to a $10,000 budget split across twenty variants, in one chart
- The four conditions under which this whole calculation stops being true
- Why we still build more than we can measure, and what that is actually for
Four figures this post rests on
The numbers under the argument01How many creatives should you test on Meta?
As many as your conversions will pay for, which for most accounts under $20,000 a month is between three and eight. The number is a function of your budget and your conversion rate, not of your ambition. A creative that has not bought enough conversions to separate itself from noise has told you nothing, however confident the dashboard looks.
Nobody publishes it this way. The pages ranking for this question assert a figure and move on: five per ad set, three to five a week, twenty a month. We went looking for an official platform number too. Meta's own Ads guide, checked on the day this was published, specifies formats, aspect ratios and file sizes, and states no recommended number of creatives at all.
What different budgets can actually buy you
Two different questions| Dimension | Can learn from | Worth shipping anyway |
|---|---|---|
| $2,000 a month, $40 CPA | No. Under 1 | 3 to 4 |
| $10,000 a month, $40 CPA | Partly. About 4 | 8 to 12 |
| $50,000 a month, $40 CPA | Yes. About 19 | 30 or more |
| $10,000 a month, $150 CPA | No. About 1 | 3 to 4 |
| $10,000 a month, $12 CPA | Yes. About 13 | 20 or more |
Four numbers the ranking pages assert, and what each one leaves out
Flip them02A creative learns from conversions, not from impressions
This is where the rules of thumb go wrong. Impressions are cheap and they feel like data, so a variant with 40,000 impressions and two purchases reads as tested. It is not tested. The thing you are trying to measure is a difference in a rate, and a rate is only as certain as the number of events underneath it.
What your money is actually buying, in order
The chainThe standard sample-size formula for comparing two rates, at 95% confidence and 80% power, simplifies neatly when your conversion rate is small. The impressions cancel and you are left with a number of conversions per variant, which is the only form of the answer that is any use to a person planning a month of spend.
The whole calculation, in five lines
Copy itconversions this month = monthly spend / your CPA
conversions per variant = 16 / (relative difference you want to catch)^2
a 50% difference -> 64 conversions
a 20% difference -> 400 conversions
creatives you can test = conversions this month / conversions per variant
spend each one needs = conversions per variant x your CPA
Worked, at $10,000 a month and a $40 CPA:
250 conversions / 64 = 3.9 creatives
and each of those needs about $2,560 before its number means anything.Two honest caveats sit on that formula. It assumes each variant gets a fair share of the spend, which on an auction platform it does not. And it assumes you decide the sample size in advance, which almost nobody does. Both make the real number worse than the arithmetic says, never better.
03What happens when you split a small budget across twenty variants
Nothing dramatic. That is the problem. Twenty variants on $10,000 a month at a $40 CPA is twelve or thirteen conversions each, and at that volume the ranking you see is mostly the order in which luck arrived. You will still pick a winner off it, because the dashboard sorts the column for you.
Conversions each variant gets, against what a verdict costs
$10,000 a month, $40 CPASee the numbers as a table
| Conversions per creative | Conversions each one gets | Conversions needed to call a 50% difference |
|---|---|---|
| 2 creatives | 125 | 64 |
| 4 creatives | 63 | 64 |
| 8 creatives | 31 | 64 |
| 20 creatives | 13 | 64 |
| 40 creatives | 6 | 64 |
The failure mode is not that you learn nothing. It is that you learn something false and then act on it with real money, which is the subject of why twenty creatives usually produce noise rather than a winner. Splitting further feels like more information and is less.
A creative that has not bought enough conversions to separate itself from noise has told you nothing, however confident the dashboard looks.
The rule this whole post is an argument for
How many new creatives did you put live last month?
The percentages here are an illustrative distribution, not survey data. The reason to vote is the comparison you are about to make: whichever bucket you land in, run your own conversions through the calculator below and see how many of those were tests and how many were shipping.
04Work out your own number
Put your four numbers in. The last output is a sanity check rather than a target: if the CPA implied by your CPM and conversion rate sits far above the CPA you are budgeting to, the creative count above is optimistic and you should plan against the implied figure.
How many creatives can you actually learn from?
Put your numbers in05Where does the arithmetic stop working?
In four places, and three of them make your real number smaller than the calculator says. Any page that gives you a creative count without naming these is selling confidence rather than method.
The four conditions that break the calculation
One per tabThe platform is not running your experiment for you
Delivery is not a randomized trial. The system concentrates budget on whatever looks promising early, so your variants do not receive equal spend and the ones that lose early may never accumulate enough events to be judged at all.
The practical consequence: your effective number of tested creatives is lower than the number you uploaded, sometimes much lower.
- Plan against the spend a variant will actually receive, not the spend you divided by hand.
- If you need a fair comparison, run a structured split test rather than reading the delivery report.
Five variations of one concept is one swing rendered five times
We learned this by retiring eight finished ads in a single afternoon. They differed on cut rhythm, transitions, type and music. To a viewer they were identical on argument, subject, world, energy, point of view and borrowed genre.
Earlier the same disease had produced twenty ads off one skeleton. The verdict written down at the time was three words: one ad, twenty times.
- Count concepts, not files. Two concepts times four executions is two tests, not eight.
- Our rule now: no two pieces in a set may match on more than two of those six axes.
You will look early, and looking early is the expensive part
Evan Miller's demonstration is the clearest one published: stop a test the moment it crosses significance and your false positive rate rises from the 5% you think you are running to about 26.1%.
He is blunt about the fix. Decide the sample size in advance, and wait.
- What you think you are running
- 5% false positives
- What stopping early gets you
- About 26.1%
- Peek ten times
- 1% significance is really 5%
The ground moves while you are measuring it
A creative's performance is not a fixed property. It decays with exposure, so a long test and a short test are measuring different objects, and a winner from six weeks ago may simply be tired rather than wrong.
This is the one condition that argues for more volume rather than less: you need replacements in the pipeline before you need them on the account. Telling decay apart from a bad week is its own problem, covered in creative fatigue or a bad week.
Four of ours. Count the swings, not the files.
Ours, and the point is uncomfortable06So should you make more than you can measure?
Yes, and the reason is not that measurement is overrated. It is that the hit rate does not improve with skill. Motion's read across 578,750 Meta creatives is that roughly 5% qualify as winners, and that top accounts and average accounts sit at about the same rate. What separates them is swings taken: 18 new ads a week surfaces roughly 0.9 winners a week, 4 a week surfaces roughly 0.2.
So volume is real leverage, and it is leverage on the production side rather than the measurement side. Which is why the shape of our own production is a surplus that gets culled, not a single idea that gets refined. The last creative act is selection.
Rework means picking, not re-directing.
The rule our production runs on
Four survivors from sets that were larger than this
What a surplus looks likeTwo things follow. Ship above your measurement line if you can afford to, but stop calling the extras a test. And make the extras genuinely distinct, because volume made of one idea inherits that idea's fate. The production economics of doing that without the cost getting away from you are in what failed renders actually cost.
07What we cannot tell you
The one thing we cannot give you is the thing a creative volume plan most wants: evidence that our creative wins more often than somebody else's. No ad in our own corpus carries a return on ad spend, so we have not measured it and neither has anyone in this category who claims it. The full version of that admission sits with the citation audit in where hook rate and hold rate benchmarks actually came from.
Before you commit next month's creative plan
Tick as you go - it remembersThat limit is worth stating plainly because it changes what this post is for. The arithmetic above is not ours and is not in dispute: it is the standard sample-size calculation, and you can check it against any calculator you like. What we contribute is the production side, which is what it costs to put a genuinely different swing in front of the auction rather than another render of the same idea. Those are separate skills and only one of them is a claim about performance.
Questions people actually ask
Open what you needHow many creatives should I test on Meta?
Divide your expected monthly conversions by about 64. At $10,000 a month and a $40 CPA that is 250 conversions and roughly four creatives you can genuinely separate from each other. Below that you are ranking noise. You may ship more than four for coverage and freshness, and that is a defensible thing to do, but it is not a test.
How many creatives per ad set should I run?
Meta's own Ads guide states no number, which is worth knowing before you take one from a blog. The practical constraint is that every additional ad in a set divides the same conversion pool, and the auction will not divide it evenly. Two to four distinct concepts per ad set is a reasonable working shape for accounts under $20,000 a month.
How many new creatives should I test a month?
Testing and shipping are separate budgets. Test as many as your conversions can carry, which is usually three to eight. Ship as many as your production can carry, because winner rates are roughly flat across accounts and volume is the lever that moves winner counts. Just do not read a verdict off the ones you shipped for volume.
Is it better to test many small variations or a few big swings?
Big swings, until you have the conversion volume to resolve small differences. A 50% difference needs about 64 conversions per variant to detect. A 20% difference needs about 400. Most accounts cannot afford the second, so testing button colors of a creative concept spends money to produce a number nobody should act on.
Does creative diversity actually matter, or is it a platform talking point?
It matters for a reason that has nothing to do with the algorithm: your effective test count is the number of genuinely different ideas you ran, not the number of files you uploaded. We retired eight finished ads at once after realizing they were identical to a viewer on argument, subject, world, energy, point of view and borrowed genre. Eight uploads, one swing.
What if my conversion volume is too low to test anything?
Then move the decision earlier, where it is cheap. Judge creative on paper and at feed size before it runs, use a higher-frequency proxy such as hook rate to catch the truly dead openings, and accept that the in-market comparison will be qualitative for now. The numbers behind those proxies deserve their own scrutiny, which is in hook rate and hold rate.
For most accounts the honest number of creatives you can learn from in a month fits on one hand. The twenty in your plan were one decision, made once, rendered twenty times, and then ranked by a column that had no right to be sorted.
Where the numbers came from
- Motion. Creative Benchmarks 2026: winners are rare (578,750 creatives, 6,015 advertiser accounts, $1.29bn Meta spend) - published 17 April 2026; a winner is defined there as a creative spending at least 10x its account median and at least $500 in total
- Evan Miller. How Not To Run an A/B Test - the 26.1% false positive rate from stopping a test the moment it looks significant
- Evan Miller. Sample size calculator for A/B tests - used as a cross-check on the approximation in this post
- Meta. Facebook Ads guide: ad format specs and recommendations - checked on 1 September 2026; it specifies formats and file specs and states no recommended number of creatives
Every figure above links to the place it was published. Numbers marked as ours are measured inside this studio and we say so where they appear. We do not print a statistic we cannot point at.
Get one creative you did not pay for, and put it in the test.
Send a link to your product and we will build one finished ad from it, free, inside three days, yours to run either way. Give it the spend the calculator above says one creative needs, and you will have a real answer instead of a ranked list.
Replies within a day. Ad within three.