One visual language per brand, and why AI defaults to the average
A generator is trained toward the middle of everything it has seen, so its default output is the average of your category. The average is nobody's brand. Here is what a visual language is actually made of, as a list of decisions you can run on your own.
What is in here
- What is a brand visual language?
- Why do all AI ads look the same?
- The eight decisions a visual language is made of
- How do you test whether you actually have one?
- Borrowing somebody else's grammar
- One owner, three properties, three languages
- Sameness is the failure a maker cannot see
- What we get wrong, and what we cannot prove
Crop the logo off three of your ads and show them to a stranger. If they cannot tell which two came from the same brand, you do not have a visual language. What you have is a list of decisions nobody made: register, camera position, grade, grain, type, transition grammar, sound, and the things the brand never does. A generator will not make them for you, because it is trained toward the middle of everything it has seen, and the middle of a category is nobody's brand.
- The eight decisions a visual language is actually made of, as a list you can run on your own brand this afternoon
- Why generated work drifts to the category average, and the one place in a brief that stops it
- The crop-the-logo test, which takes four minutes and three ads
- The device we banned from every client brand, and what it taught us about borrowing somebody else's grammar
- Why one owner with three properties got three languages instead of one, and what merging them would have cost
Four numbers from our own sameness problem
Measured in this studio01What is a brand visual language?
It is the set of decisions that make your ads recognizable before anyone reads a word. Not colors and a logo lockup. Decisions: where the camera stands, how light behaves, what the type weighs, how one shot becomes the next, what the ears get, and what you refuse to do. A style guide describes. A language decides.
Style guide against visual language
Two documents| Dimension | Brand style guide | Visual language |
|---|---|---|
| What it contains | Hex values, clear space, logo misuse, approved fonts | Yes. Decisions about shots, light, pace, type weight and sound |
| Question it answers | Is this on brand? | Yes. What do we shoot on Tuesday, and how do we cut it? |
| Survives the logo being cropped off | No. No. Remove the mark and the document has nothing left to say | Yes. That is the whole test |
| Usable as a generation brief | No. Barely. A palette is not a prompt | Directly. Every decision is a token that does work |
| Contains a list of refusals | No. Rarely, and never about behavior | Always. The never-list is the fastest part to build |
| Length | 40 pages | Yes. One page, and it changes what gets made |
The reason this distinction earns its keep right now is that a style guide cannot constrain a generator in any useful way. You can hand a model your hex codes and it will honor them while producing a shot that any competitor in your category could have run. Color compliance is not identity. What a model needs is a decision about behavior, and behavior is exactly what a guide leaves out.
02Why do all AI ads look the same?
Because a generator produces a likely sample from everything it has learned, and the likeliest image in any category is the average of that category. Ask for a skincare ad and you get the skincare ad: soft window light from the left, a marble surface, a dewy drop, a serif in the corner. It is competent, it is nobody's, and your competitor gets the same one. We wrote that one out shot by shot in the beauty ad everyone has already seen, along with what a brand in that category can do instead.
Averaging is the feature, not the bug
This is not a flaw to be fixed by a better model. It is what the machine is for. Averaging is the useful behavior when you ask for a coffee cup and want a coffee cup. It becomes a problem the moment you want a coffee cup that could only belong to one brand. Every distinguishing decision has to be supplied by you, in the prompt, in the assets, in the cut, or it defaults to the middle. The sibling problem, which bites in longer sets, is per-render drift between shots.
The same product, five decisions away from the average
Build it from partsA generator gives you the average of a category. The average is nobody's brand.
the sentence we open every brand kit with
03The eight decisions a visual language is made of
Run this on your own brand. Most of it you already decide by accident, one ad at a time, which is why it drifts. Writing the eight down takes an afternoon and it converts a taste you hold into an instruction somebody else, or a model, can follow.
Write these eight, then hand them to whoever makes your next ad
Tick as you go - it remembersFour of the eight, in detail
Switch between themWhere the camera stands is your loudest cue
Viewers cannot name a focal length and they recognize one instantly. A brand that always shoots from slightly above, or always sits at table height, is legible before the product is.
- Write it as a position, not an adjective: overhead at 90 degrees, hands enter from the bottom
- The cheapest production upgrade is not a better camera. It is putting the camera somewhere a camera does not normally go
- Generated shots inherit the average camera unless the prompt names the position
The two that make mixed sources read as one
Most short-form ads now carry three or four kinds of footage: real product photography, phone footage, stock, generated coverage. The grade and one grain plate are what stop that being visible.
- One grain plate, tiled, never scaled, re-seated every frame
- A grade is a rule about what you will not touch as much as a look
- If a viewer can tell which shot came from where, you do not have a language yet, you have a folder
Weight, not size, decides whether it reads
An ultralight face at 36px integrates to a stem 0.71 pixels wide. No stroke ever fully covers a pixel column, and only six or seven percent of the ink reaches full opacity. At regular weight the same stem is 3.70 pixels and 78 percent of the ink is solid.
A size bump does not fix a sub-pixel stem. Only weight does.
- Our cap-height floor
- At least 5% of frame height, measured off the plate's own ink
- Our dwell floor
- 1.2 seconds, or 0.35 seconds per word, whichever is longer
The half of the language nobody writes down
Tempo is a per-brand and per-ad decision. Across twenty-one reference films we measured tempos from 65 to 144 BPM. When one of our own sets put twenty ads on a single 90 BPM grid, the result was audible as sameness regardless of instrument.
The sound world belongs in the language document next to the grade, and it gets decided on paper before a frame is cut.
- Bass-led or voice-led, never neither
- Something on the soundtrack at frame zero, never a fade in
- A great bed will not rescue a dead opening frame. A bad one loses a viewer the picture already caught
One warning about the list. Cutting rhythm plus light over static compositions is itself a language, and it is the one most people miss, because it does not look like design. In our own proven set the elements barely move at all. What moves is the light patch traveling across a plate, the crop window walking, and the cut. If your brand wants events on screen rather than light, read when a declared style beats a bad photoreal before you brief them.
Where the brand mark actually goes
Second by secondRead it as a list
| At | Channel | What happens |
|---|---|---|
| 0.0s | Picture | Hook, resolving |
| 3.2s | Picture | Body: the argument, in the brand's own grammar |
| 9.5s | Picture | Payoff and end card |
| 3.2s | Brand mark | First appearance, held 1s at full opacity |
| 13.4s | Brand mark | Repeat, not an introduction |
04How do you test whether you actually have one?
The test is cheap because it removes the only cue most brands are actually leaning on. A wordmark is recognition you rented. Everything else in the frame is recognition you own, or do not. Four minutes, no research budget, and a stranger's shrug is a finding.
The crop-the-logo test, run properly
One screen at a timeFour frames, four languages, no wordmarks
Ours, logos cropped05Borrowing somebody else's grammar
We have a device in the house that is banned on every client brand. It is a full-frame color field held for seven or eight frames, fired between shots. In the brand it came from it means something: that brand has nine products and each one owns a signature color, so a field cutting to a product's own color is a sentence about which product you are looking at.
We ported it into a client engine as craft. First on a timer, then, after being caught, only at chapter boundaries, which was still keeping the device. It read as a slide transition. Every film in that set was re-cut to remove it, and the build engine now refuses that shot type with an error that names the reason, because habit had already brought it back twice.
The question that would have caught it in ten seconds
The terms we use for this, defined
Search the vocabularyVisual languageBrand
The six perceptual axesSets
Brand-reveal windowTiming
Flash fieldBanned device
Legibility floorType
Margin ruleFootage
06One owner, three properties, three languages
Our hotel client owns three properties. The cheap answer, the one we were tempted by, was a single house language across all three: one grade, one type system, one music bed, three logos. It would have halved the work. We treat them instead as three separate brands, with three languages and a hard rule that a creative may only use one property's own photography.
Merge them, or give each one a language
The decision, filled in honestly| Dimension | One house language | Three languages |
|---|---|---|
| Production cost | Yes. Roughly half. One kit, one bed, one grade | Three of everything, three corpus studies |
| What a returning guest recognizes | No. The group. Which is not the thing they book | The property they liked, in the first second |
| When two properties advertise the same week | No. They compete against each other looking identical | They read as alternatives, which is what they are |
| Risk of a shot from the wrong property | No. High, and invisible until a guest arrives | Structurally impossible. One property per creative, enforced |
| Speed to the first cut | Yes. Faster. Everything is already decided | Slower on brand one, then normal |
| What happens if a property is sold | No. Its ads take the group's identity with them | It leaves with its own language, intact |
The enforcement matters more than the intention. A rule that lives in a document gets broken by whoever is tired at 2am. The one-property-per-creative rule lives in code, and a build that mixes two properties refuses to run. There is a longer version of this argument in three hotels, three languages, including what we got wrong on the first pass.
Three languages, running
Ours, three brands07Sameness is the failure a maker cannot see
Eight of our finished ads were retired in one go. They differed on cut rhythm, on transitions, on type and on music, and to a viewer they were identical on argument, subject, world, energy, point of view and borrowed genre. We had been measuring the wrong things: editing mechanics, which vary easily, instead of perception, which does not.
Ninety-four images, nine camera angles. The same disease showing up in a second organ.
our own audit, after the same problem appeared in statics
Five things people believe about brand consistency
Flip them08What we get wrong, and what we cannot prove
The first honest thing: nothing in our own corpus carries a click-through rate or a return on spend, so the case for a language rests on recognition and on the cost of being interchangeable rather than on a measured lift. Anyone quoting you a number for this should be asked where they got it.
A language and a fingerprint are the same thing from two sides
Hold a set of decisions hard enough and the work starts repeating itself. We now cap repeated phrasing across a set and require a rare device to stay rare, because running an unusual move at eight times its natural frequency destroys the thing that made it worth having. The sound half of the language has its own post, sound is 40% of the ad, and it is the half most brands have never written down at all.
The last time you briefed an ad, what did you hand over as the brand's look?
These percentages are an illustrative split, not survey data. The pattern we see in real briefs is that the last option is rare, and it is the only one of the four a generator can actually act on.
Questions people actually ask
Open what you needWhat is the difference between a visual language and a style guide?
A style guide describes what the brand looks like when it is standing still: colors, fonts, clear space, logo misuse. A visual language decides how the brand behaves in motion: where the camera stands, how light falls, what the type weighs, how a shot ends, what the ears get, and what the brand refuses to do. The guide protects the mark. The language is what a stranger recognizes with the mark cropped off.
Why do all my AI-generated ads look the same as everyone else's?
Because a generator returns a likely sample from what it has learned, and the likeliest image in a category is that category's average. Soft window light, marble, a dewy drop, a serif in the corner.
The fix is not a better model or a longer prompt. It is supplying the distinguishing decisions yourself: camera position, light behavior, grade, grain, type weight and sound. Anything you do not decide, the model decides toward the middle.
How do I build a visual identity for ads if I only have phone photos?
Start with the negative list and the camera position, because both are free. Decide three things you will never do, and decide exactly where the camera lives. Then pick one light condition you can actually get repeatedly, which for most small brands is one window at one time of day. Three decisions held for six ads will out-recognize a forty-page guide that nobody applies.
How many ads before a visual language starts working?
Recognition needs repetition, so the honest answer is more than you would like and fewer than a rebrand. What we can say from our own work is that the decisions have to be written before the ads, not discovered after them. A language derived retrospectively from six ads you already made will mostly encode the accidents.
Can two brands in the same category share a visual language?
They can, and one of them is wasting money. If a viewer cannot separate you at 360 pixels with the logo cropped, your media spend is partly buying awareness for whoever else looks like you. The cheapest correction is not a new palette. It is a different camera position and a different sound world, both of which are free.
Does a visual language limit what we can test?
It limits the wrong axis and frees the right one. Testing five transitions of one concept is one swing rendered five times. A language fixes the surface so that your tests can vary the things that actually move results: the argument, the subject, the world, the energy, the point of view and the genre you are borrowing.
A template is a machine for producing the average, and the average is available to your competitor for free, this afternoon, from one prompt.
Where the numbers came from
- arXiv. Ho, Jain and Abbeel, Denoising Diffusion Probabilistic Models - for what a generator is actually doing when it produces a sample
- Ehrenberg-Bass Institute for Marketing Science. Institute research on distinctive brand assets - the body of work behind the idea that brand codes, not logos, do the recognizing
- W3C. Understanding Success Criterion 1.4.3: Contrast (Minimum) - the 4.5:1 ratio our type floors are built on
Every figure above links to the place it was published. Numbers marked as ours are measured inside this studio and we say so where they appear. We do not print a statistic we cannot point at.
Send a link. Get one finished ad back.
One finished cut built from your own product, inside three days, free and yours to run whether or not we ever work together. It arrives with the eight decisions written down, so you can keep the language even if you never work with us again.
Replies within a day. Ad within three.