Scale anchoring: why your AI ad gets the size of things wrong
Generated footage has no unit of measurement, so it renders your product at the average size of everything it has ever seen. Here is the failure nobody names, the kill that taught us to write it down, and the one line that fixes it in a brief.
What is in here
A generator has no unit of measurement. It has never held your bottle, so it renders size from whatever its training pictures implied, and the implication skews large because advertising photography skews large. No adjective fixes this. An anchor does: a known object in frame at a known size, the landmark's real measured dimension written beside the height of the person standing next to it, and the relationship between the two said out loud in the shot line. Scale is briefed. It is never hoped for.
- The five cues a human eye reads size from, and which of them a generator actually controls
- The kill that made us write scale into every brief: a stone crest that rendered about eight feet tall
- Why reads shoulder-height beside her generates better than 5ft
- A sixty-second review pass that catches a scale error before a client does
Everything in this frame that tells you how big the bottle is
Tap the numbers
01Why does AI get the size of things wrong?
Because there is no metric layer anywhere in the process. A model samples pixels that look plausible next to each other. It does not hold a tape. Ask for a stone crest and you get the average of every stone crest in the training set, and that average is whatever the photographers who shot those crests happened to be standing near.
The error would be survivable if it were random. It is not. Generated objects skew big, because the pictures models learned from skew big: products get shot low and close so they look important, buildings get shot to look imposing. A model trained on advertising inherits advertising's flattery along with its lighting.
What your eye is doing that the model is not
You judge size from five things at once, and none of them is the object itself. A familiar item in frame, borrowed as a ruler. The horizon, which tells you where the camera is standing. Perspective compression, which is really you reading a focal length. Depth of field, since only a close lens throws a near edge soft. And contact, the small dark seam where a thing meets the surface it is resting on. A generator reproduces the look of all five and controls none of them.
The five cues your viewer reads size from
Who controls what| Dimension | Left to the model | Written into the shot line |
|---|---|---|
| A familiar object in frame | No. Rendered at the training-set average | Yes. Name the object and its real size |
| The horizon and camera height | No. Usually absent, sometimes invented | Yes. State the lens height: seated, standing, floor |
| Perspective compression | Partly. Follows the look, not a lens | Yes. Name a focal length. Models answer 85mm better than compressed |
| Depth of field | Partly. Copies the style, ignores the optics | Yes. Say what is sharp and what is not, in distances |
| Contact and shadow | No. Drawn where it looks pleasant | Yes. Say what it rests on and how heavy it is |
| A person for reference | No. Sized to suit the composition | Yes. Give a height, then relate the object to it |
02The crest that rendered about eight feet tall
A hotel client's entry crest is a stone piece by the door. It measures roughly five feet. In a film we built, it rendered at about eight, taller than the creator walking past it, and the shot went through three separate reviews before anybody said anything. Everyone was looking at the light. The light was good.
The owner spotted it in under a second, because he walks past the real thing every day of his life. That is the part worth sitting with. This was not a hallucinated palace or an invented lobby. Every surface in the shot was a real place we had photographs of.
The space was entirely real, and the film still lied.
Our own note on the crest kill
What that single shot cost
Counted off one shotThe second half of the rule is stranger and it matters more. The row carries the number for the reviewer. The prompt carries the relationship for the model, because a model cannot check a number but it can place a thing against a person.
One row of a content sheet, in the form we actually write it
Copy this shapeSHOT 04 - entry crest, creator walks past, camera left
creator height ..... 5'6" (owner-stated)
crest height ....... 5'0" (MEASURED, tape)
relationship ....... crest top sits just above her shoulder
prompt phrasing .... "the stone crest reads shoulder-height
beside her as she passes it"
camera ............. 85mm equivalent, lens at 5'2", static
provenance ......... source photo, west elevationThe provenance tag is doing more work than it looks like. Measured means somebody put a tape on it. Owner-stated means somebody who lives there said so. ASSUMED, in capitals, means we guessed, and a reviewer reading that row knows the shot is standing on a guess. The eight-foot crest had no row at all, which is a worse state than a marked guess.
03How do you write scale into a brief so it survives the render?
Write the number and the relationship, in that order, on the shot's own line. The number gives your reviewer something to check against. The relationship gives the model something it can act on, because generators respond to spatial relationships between objects far better than they respond to units. Five feet means nothing to a sampler. Shoulder-height beside a woman means a great deal.
Build the scale line
Assemble it04Product scale: a bottle, a patch, a ring, a mattress
The hotel version of this failure is dramatic and easy to tell as a story. The product version is quieter and costs more, because a slightly oversized serum bottle does not look wrong. It looks like a different, cheaper product. Scale sits beside the other things current models still break, which we catalog in the physics test. Here is what each category gets wrong, and what anchors it.
Four products, four anchors, four tells
Pick a categoryThe apothecary drift
Ask for a 30ml dropper bottle and you often get something closer to a 200ml apothecary jar with a dropper on top. The proportions inside the object are what give it away: on a real 30ml bottle the glass pipette runs most of the bottle's height, and the cap is wide relative to the shoulder.
- Anchor: a hand, or a second object with a fixed size such as a folded cloth
- Tell: cap width against bottle width. A small bottle has a proportionally large cap
- Best anchor
- Fingers wrapped around the glass
- Worst frame
- Bottle alone on an empty surface
The coaster problem
A patch is a thin disc, and a thin disc has no intrinsic size. In our own work on Ephoria, our house wellness-patch brand, generated patch shots came back at anything from a coin to a coaster inside the same batch, with nothing in frame to settle the argument.
- Anchor: a fingertip lifting it, or the pouch it came out of
- Tell: how much of the palm it covers when it is being applied
- Best anchor
- The moment of application, on skin
- Worst frame
- The patch photographed flat, alone, from directly above
Where jewelry turns into costume
Rings fail on band thickness rather than on overall size. A generated band tends to run heavy, the stone runs proportionally large, and the result reads as costume rather than fine jewelry even when nothing is technically wrong. Put a hand in the shot and it gets worse before it gets better, which is the contact problem in AI UGC.
- Anchor: the knuckle, and the finger next to the one wearing it
- Tell: band thickness against the width of the finger under it
- Best anchor
- Worn, with the adjacent finger visible
- Worst frame
- The ring on a plinth with no hand anywhere
One size class down
Big soft objects drift toward whatever fills the frame nicely. A king renders like a double, or a double renders like a king, and the giveaway is almost never the bed. It is the pillows, which the model sizes to the bed rather than sizing the bed to them.
- Anchor: standard pillows, a door frame, a bedside lamp
- Tell: how many pillow widths fit across the head of the bed
- Best anchor
- A person sitting on the edge, feet on the floor
- Worst frame
- Bed shot square on with cropped walls

05How do you catch a scale error in under a minute?
Watch the cut once at phone size with the sound off and ask one question per shot: what in this frame tells me how big that is? If the answer is nothing, you have found the risk whether or not the shot is wrong. Then check the two frames on either side of every cut, because scale errors hide in the join.
The sixty-second scale pass
Tick as you go - it remembers06What anchoring does not fix
Getting size right is a floor, not a lever. It stops an ad from being wrong. It does not make an ad work. The closest measured number anyone has put on craft of this kind is CreativeX's: about ten percent more Creative Quality Score buys around two percent off CPM, across roughly 822,000 observations. Real, solid, and small. Anyone selling you scale accuracy as a growth strategy is selling you hygiene at a premium.
Four things we have been told about scale, and what we found
Flip themYour audience is more suspicious than your client thinks
The IAB put 505 consumers and 104 executives on the same question and found advertisers overestimating consumer warmth toward AI ads by 37 points: 82 percent of executives against 45 percent of consumers. A viewer who has already decided a frame is machine-made starts hunting for proof, and size is the cheapest proof available. It takes no expertise at all. Everybody knows how big a door is.
When a generated shot feels off but you cannot say why, what do you look at first?
Most people go to texture, because texture is what the discourse talks about. In our own review notes the thing that actually turns out to be wrong is far more often a relationship between two objects: how big, how far, how heavy.
Anchoring also cannot save you from a shot the space never supported. If nobody has photographed the thing from that angle, we refuse the move rather than invent it, and we name the one frame that would unlock it. A refusal like that turns a no into a ten-minute shoot list, which is the argument running underneath our piece on hotel ad creative that does not lie about the property.
Questions people actually ask
Open what you needWhy does AI get the size of objects wrong?
Because nothing in the generation process measures anything. The model produces pixels that are statistically plausible together, and size is inferred from the composition rather than imposed on it. The bias runs large, because commercial photography shoots products low and close and shoots buildings to look imposing, and a model trained on that inherits the flattery.
How do I make an AI video get proportions right?
Put a second object of known size in the frame and describe the relationship between the two, rather than stating a measurement. A hand, a door, a phone, a pillow, or a person with a stated height all work.
Then write the real number on the shot line for your reviewer, marked as measured, stated by the owner, or assumed. The number is what makes the error checkable later.
Why does my product look bigger in the AI render than in real life?
Two causes stack. The training data leans toward hero product photography, which is shot to make things look substantial. And your frame probably has no anchor, so nothing contradicts the model's default. Add a hand or a known object, name a longer focal length, and state the camera height.
Can I just fix scale in post?
Sometimes, and it is usually the wrong trade. Rescaling a generated object breaks its contact shadow and its depth of field at the same time, so you fix one tell and create two. Our rule on generated footage is that a bad take gets re-rolled rather than repaired, and scale is the clearest case for it.
Does this apply to still ad creative as well as video?
Yes, and it is harder to catch, because a still gives the eye time to settle on a wrong size instead of racing past it. The one advantage is that a static frame lets you compose an anchor deliberately: a hand entering the frame, a folded cloth, the edge of a table.
How do you check scale on a property you have never visited?
You ask, and you write down which answers were measured and which were guessed. For our hotel work every landmark carries its real dimension with the provenance marked, and an assumed dimension is flagged so that a reviewer knows the shot is resting on a guess rather than on a tape measure.
Our automated gates never caught the crest, and they still would not. What caught it was a person who knew the real thing. Build your review around that fact, and write down every dimension a stranger would otherwise have to guess at.
Where the numbers came from
- IAB. The AI Gap Widens - 505 consumers and 104 executives; used for the 37-point perception gap
- CreativeX. Creative Quality Score - used for the quality-score to CPM relationship across roughly 822,000 observations
- Motion. Creative Benchmarks 2026: winners are rare - 550,000+ ads, 6,000+ advertisers, about $1.3bn in spend
Every figure above links to the place it was published. Numbers marked as ours are measured inside this studio and we say so where they appear. We do not print a statistic we cannot point at.
Send a link. Get one finished ad back.
One finished cut built from your own product, inside three days, free and yours to run whether or not we ever work together. Check every frame of it for an anchor. If you find a shot where nothing tells you how big your product is, we broke the rule in this piece.
Replies within a day. Ad within three.