Sutra

One visual language per brand, and why AI defaults to the average

A generator is trained toward the middle of everything it has seen, so its default output is the average of your category. The average is nobody's brand. Here is what a visual language is actually made of, as a list of decisions you can run on your own.

What is in here
  1. What is a brand visual language?
  2. Why do all AI ads look the same?
  3. The eight decisions a visual language is made of
  4. How do you test whether you actually have one?
  5. Borrowing somebody else's grammar
  6. One owner, three properties, three languages
  7. Sameness is the failure a maker cannot see
  8. What we get wrong, and what we cannot prove
The short answer

Crop the logo off three of your ads and show them to a stranger. If they cannot tell which two came from the same brand, you do not have a visual language. What you have is a list of decisions nobody made: register, camera position, grade, grain, type, transition grammar, sound, and the things the brand never does. A generator will not make them for you, because it is trained toward the middle of everything it has seen, and the middle of a category is nobody's brand.

What you get out of this
  1. The eight decisions a visual language is actually made of, as a list you can run on your own brand this afternoon
  2. Why generated work drifts to the category average, and the one place in a brief that stops it
  3. The crop-the-logo test, which takes four minutes and three ads
  4. The device we banned from every client brand, and what it taught us about borrowing somebody else's grammar
  5. Why one owner with three properties got three languages instead of one, and what merging them would have cost

Four numbers from our own sameness problem

Measured in this studio
8finished ads retired in one verdict: different cuts, identical argument, world and point of viewSutra Haus verdict log
20ads that turned out on inspection to be one ad, twenty timesSutra Haus verdict log
94static images that shared nine camera angles across eighty-five layoutsSutra Haus audit
6perceptual axes we now measure a set against. No two pieces may match on more than twoSutra Haus
Ours, counted off delivered work. We are not describing a mistake other studios make. All four of these are sets we built, shipped internally and then had to retire.

01What is a brand visual language?

It is the set of decisions that make your ads recognizable before anyone reads a word. Not colors and a logo lockup. Decisions: where the camera stands, how light behaves, what the type weighs, how one shot becomes the next, what the ears get, and what you refuse to do. A style guide describes. A language decides.

Style guide against visual language

Two documents
Style guide against visual language
DimensionBrand style guideVisual language
What it containsHex values, clear space, logo misuse, approved fontsYes. Decisions about shots, light, pace, type weight and sound
Question it answersIs this on brand?Yes. What do we shoot on Tuesday, and how do we cut it?
Survives the logo being cropped offNo. No. Remove the mark and the document has nothing left to sayYes. That is the whole test
Usable as a generation briefNo. Barely. A palette is not a promptDirectly. Every decision is a token that does work
Contains a list of refusalsNo. Rarely, and never about behaviorAlways. The never-list is the fastest part to build
Length40 pagesYes. One page, and it changes what gets made
You need both. The guide protects the trademark and keeps the pack shot honest. The language is what a stranger recognizes at speed, and almost nobody writes it down.

The reason this distinction earns its keep right now is that a style guide cannot constrain a generator in any useful way. You can hand a model your hex codes and it will honor them while producing a shot that any competitor in your category could have run. Color compliance is not identity. What a model needs is a decision about behavior, and behavior is exactly what a guide leaves out.

02Why do all AI ads look the same?

Because a generator produces a likely sample from everything it has learned, and the likeliest image in any category is the average of that category. Ask for a skincare ad and you get the skincare ad: soft window light from the left, a marble surface, a dewy drop, a serif in the corner. It is competent, it is nobody's, and your competitor gets the same one. We wrote that one out shot by shot in the beauty ad everyone has already seen, along with what a brand in that category can do instead.

Averaging is the feature, not the bug

This is not a flaw to be fixed by a better model. It is what the machine is for. Averaging is the useful behavior when you ask for a coffee cup and want a coffee cup. It becomes a problem the moment you want a coffee cup that could only belong to one brand. Every distinguishing decision has to be supplied by you, in the prompt, in the assets, in the cut, or it defaults to the middle. The sibling problem, which bites in longer sets, is per-render drift between shots.

The same product, five decisions away from the average

Build it from parts

Each option is a real token we use, with the reason it does work. Pick one from every row and read what you get. The point is not the sentence. It is that five decisions, made once and held, are a language.

A generator gives you the average of a category. The average is nobody's brand.

the sentence we open every brand kit with

03The eight decisions a visual language is made of

Run this on your own brand. Most of it you already decide by accident, one ad at a time, which is why it drifts. Writing the eight down takes an afternoon and it converts a taste you hold into an instruction somebody else, or a model, can follow.

Write these eight, then hand them to whoever makes your next ad

Tick as you go - it remembers
0%
Eight lines. If you can answer seven and not the eighth, the eighth is the one doing the damage, and it is usually the last one.

Four of the eight, in detail

Switch between them
Where the camera stands is your loudest cue

Viewers cannot name a focal length and they recognize one instantly. A brand that always shoots from slightly above, or always sits at table height, is legible before the product is.

  • Write it as a position, not an adjective: overhead at 90 degrees, hands enter from the bottom
  • The cheapest production upgrade is not a better camera. It is putting the camera somewhere a camera does not normally go
  • Generated shots inherit the average camera unless the prompt names the position
These four come up most often in briefs, and they are the four a generator will decide for you if you do not.

One warning about the list. Cutting rhythm plus light over static compositions is itself a language, and it is the one most people miss, because it does not look like design. In our own proven set the elements barely move at all. What moves is the light patch traveling across a plate, the crop window walking, and the cut. If your brand wants events on screen rather than light, read when a declared style beats a bad photoreal before you brief them.

Where the brand mark actually goes

Second by second
0s1s2s3s4s5s6s7s8s9s10s11s12s13s14s15sHook, resolvingBody: the argument, in the brand's own grammarPayoff and end cardRepeat, no…PICTUREBRAND MARK
Read it as a list
AtChannelWhat happens
0.0sPictureHook, resolving
3.2sPictureBody: the argument, in the brand's own grammar
9.5sPicturePayoff and end card
3.2sBrand markFirst appearance, held 1s at full opacity
13.4sBrand markRepeat, not an introduction
Our default window is a first brand appearance between 20% and 30% of runtime, held at least a second at full opacity, with the end card as a repeat rather than an introduction. It is a synthesis of two measured positions, not a law of nature. What is not negotiable is that the direction lock names the beat, so it is never accidental.

04How do you test whether you actually have one?

The test is cheap because it removes the only cue most brands are actually leaning on. A wordmark is recognition you rented. Everything else in the frame is recognition you own, or do not. Four minutes, no research budget, and a stranger's shrug is a finding.

The crop-the-logo test, run properly

One screen at a time
Step 1
Take three of your ads and two competitors'

Screenshot a mid-film frame from each. Not the end card, which is where the logo lives and where every brand looks like itself.

Five images total. Two of your three should be from different campaigns, ideally months apart.

Step 2
Crop out every wordmark and pack shot

Also crop the product if it is distinctive enough to answer the question by itself. You are testing the language, not the packaging.

Resize everything to 360 pixels wide, which is roughly the width a feed tile actually renders at on a phone.

Step 3
Ask one question, of somebody outside the company

Which two of these came from the same brand? Do not explain the test first. Do not show them your guide.

Note how long they take. Hesitation is data.

Step 4
Ask them why, and write down the words they use

The reasons matter more than the answer. If they say the light, the pace or the type, those are your real distinctive assets and you should protect them.

If they say the color, you have one asset and it is the one easiest for a competitor to take.

Step 5
Run it again with the sound on, picture hidden

Play three ads with the screen turned away. Same question. Most brands fail this one completely, which is the cheapest opportunity in the whole exercise.

A sound world is a brand asset and almost nobody has written one down.

1 / 5

Run it cold and run it once. The second viewing is worthless because the person has already learned the answer.

Four frames, four languages, no wordmarks

Ours, logos cropped
Frames from four of our own films with the branding out of shot. Hard window light and a single gesture. Architecture framing a person kept small. A face at macro on saturated seamless. A daylight macro of a process. None of these could be swapped into another of the four.

05Borrowing somebody else's grammar

We have a device in the house that is banned on every client brand. It is a full-frame color field held for seven or eight frames, fired between shots. In the brand it came from it means something: that brand has nine products and each one owns a signature color, so a field cutting to a product's own color is a sentence about which product you are looking at.

We ported it into a client engine as craft. First on a timer, then, after being caught, only at chapter boundaries, which was still keeping the device. It read as a slide transition. Every film in that set was re-cut to remove it, and the build engine now refuses that shot type with an error that names the reason, because habit had already brought it back twice.

The question that would have caught it in ten seconds

The rule that came out of itBefore you port a move from a reference you admire, ask whether it is portable craft or that brand's own grammar. A device that carries meaning inside one brand's system carries nothing outside it, and the audience reads the emptiness as a template. The check takes one sentence: what does this mean here, and who told the viewer?

The terms we use for this, defined

Search the vocabulary
6 terms
Visual languageBrand
The set of behavioral decisions that make a brand's work recognizable with the logo cropped off: register, camera position, grade, grain, type system, transition grammar, sound world, and the list of refusals.
The six perceptual axesSets
Argument, subject, world, energy, point of view and borrowed genre. No two pieces in one set may match on more than two of them. It replaced a check that measured editing mechanics and missed sameness entirely.
Brand-reveal windowTiming
The span of runtime in which the brand mark first appears, defaulting to 20 to 30 percent of the film, after the hook has resolved. Never accidental: the direction lock names the beat it lands on.
Flash fieldBanned device
A full-frame color field held for seven or eight frames as a transition. It belongs to one brand's nine-color product system and is banned on every client brand, because elsewhere it reads as a slide transition.
Legibility floorType
A two-part gate measured on the delivered file: cap height as a fraction of frame height, and perceptual contrast against the worst local background under each glyph run, checked at delivery width and again at 360 pixels.
Margin ruleFootage
A client's own footage is never used raw. It is recropped, regraded and recomposed, and their captions and phone audio never survive. Borrowed material is punctuation in a film, never its spine.
House vocabulary, written so you can lift it straight into your own brief. Where a floor has an external standard behind it, the standard is linked in the sources at the bottom of the page.

06One owner, three properties, three languages

Our hotel client owns three properties. The cheap answer, the one we were tempted by, was a single house language across all three: one grade, one type system, one music bed, three logos. It would have halved the work. We treat them instead as three separate brands, with three languages and a hard rule that a creative may only use one property's own photography.

Merge them, or give each one a language

The decision, filled in honestly
Merge them, or give each one a language
DimensionOne house languageThree languages
Production costYes. Roughly half. One kit, one bed, one gradeThree of everything, three corpus studies
What a returning guest recognizesNo. The group. Which is not the thing they bookThe property they liked, in the first second
When two properties advertise the same weekNo. They compete against each other looking identicalThey read as alternatives, which is what they are
Risk of a shot from the wrong propertyNo. High, and invisible until a guest arrivesStructurally impossible. One property per creative, enforced
Speed to the first cutYes. Faster. Everything is already decidedSlower on brand one, then normal
What happens if a property is soldNo. Its ads take the group's identity with themIt leaves with its own language, intact
We do not win every row. Merging is genuinely cheaper and genuinely faster, and for a group selling one undifferentiated product it would be the right call. It was the wrong call here because a guest chooses between these three, not between this group and another.

The enforcement matters more than the intention. A rule that lives in a document gets broken by whoever is tired at 2am. The one-property-per-creative rule lives in code, and a build that mixes two properties refuses to run. There is a longer version of this argument in three hotels, three languages, including what we got wrong on the first pass.

Three languages, running

Ours, three brands
Ephoria - campaign ad
Hotel client - brand reel
Beauty device - campaign ad
Ephoria - brand reel
One house wellness brand, one hotel client, one beauty device client. Watch them back to back with sound. The differences that matter are pace, light and what the sound is doing, not the palette.

07Sameness is the failure a maker cannot see

Eight of our finished ads were retired in one go. They differed on cut rhythm, on transitions, on type and on music, and to a viewer they were identical on argument, subject, world, energy, point of view and borrowed genre. We had been measuring the wrong things: editing mechanics, which vary easily, instead of perception, which does not.

Ninety-four images, nine camera angles. The same disease showing up in a second organ.

our own audit, after the same problem appeared in statics

Five things people believe about brand consistency

Flip them
The fourth card is the expensive one. It is the belief that keeps a brand producing variations of one ad and calling it a test.

08What we get wrong, and what we cannot prove

The first honest thing: nothing in our own corpus carries a click-through rate or a return on spend, so the case for a language rests on recognition and on the cost of being interchangeable rather than on a measured lift. Anyone quoting you a number for this should be asked where they got it.

A language and a fingerprint are the same thing from two sides

Hold a set of decisions hard enough and the work starts repeating itself. We now cap repeated phrasing across a set and require a rare device to stay rare, because running an unusual move at eight times its natural frequency destroys the thing that made it worth having. The sound half of the language has its own post, sound is 40% of the ad, and it is the half most brands have never written down at all.

Vote, then see

The last time you briefed an ad, what did you hand over as the brand's look?

Questions people actually ask

Open what you need
What is the difference between a visual language and a style guide?

A style guide describes what the brand looks like when it is standing still: colors, fonts, clear space, logo misuse. A visual language decides how the brand behaves in motion: where the camera stands, how light falls, what the type weighs, how a shot ends, what the ears get, and what the brand refuses to do. The guide protects the mark. The language is what a stranger recognizes with the mark cropped off.

Why do all my AI-generated ads look the same as everyone else's?

Because a generator returns a likely sample from what it has learned, and the likeliest image in a category is that category's average. Soft window light, marble, a dewy drop, a serif in the corner.

The fix is not a better model or a longer prompt. It is supplying the distinguishing decisions yourself: camera position, light behavior, grade, grain, type weight and sound. Anything you do not decide, the model decides toward the middle.

How do I build a visual identity for ads if I only have phone photos?

Start with the negative list and the camera position, because both are free. Decide three things you will never do, and decide exactly where the camera lives. Then pick one light condition you can actually get repeatedly, which for most small brands is one window at one time of day. Three decisions held for six ads will out-recognize a forty-page guide that nobody applies.

How many ads before a visual language starts working?

Recognition needs repetition, so the honest answer is more than you would like and fewer than a rebrand. What we can say from our own work is that the decisions have to be written before the ads, not discovered after them. A language derived retrospectively from six ads you already made will mostly encode the accidents.

Can two brands in the same category share a visual language?

They can, and one of them is wasting money. If a viewer cannot separate you at 360 pixels with the logo cropped, your media spend is partly buying awareness for whoever else looks like you. The cheapest correction is not a new palette. It is a different camera position and a different sound world, both of which are free.

Does a visual language limit what we can test?

It limits the wrong axis and frees the right one. Testing five transitions of one concept is one swing rendered five times. A language fixes the surface so that your tests can vary the things that actually move results: the argument, the subject, the world, the energy, the point of view and the genre you are borrowing.

A template is a machine for producing the average, and the average is available to your competitor for free, this afternoon, from one prompt.

Where the numbers came from

  1. arXiv. Ho, Jain and Abbeel, Denoising Diffusion Probabilistic Models - for what a generator is actually doing when it produces a sample
  2. Ehrenberg-Bass Institute for Marketing Science. Institute research on distinctive brand assets - the body of work behind the idea that brand codes, not logos, do the recognizing
  3. W3C. Understanding Success Criterion 1.4.3: Contrast (Minimum) - the 4.5:1 ratio our type floors are built on

Every figure above links to the place it was published. Numbers marked as ours are measured inside this studio and we say so where they appear. We do not print a statistic we cannot point at.

Badal Kariwal

Runs Sutra Haus, a one-person ad studio that has shipped over a thousand finished creatives - film and stills - for DTC brands and hotels. Writes here about what the work actually taught him, including the parts that failed. The person who reads your brief is the person who builds the work. Send him something to make.

The test, run on your brand instead of ours

Send a link. Get one finished ad back.

One finished cut built from your own product, inside three days, free and yours to run whether or not we ever work together. It arrives with the eight decisions written down, so you can keep the language even if you never work with us again.

Replies within a day. Ad within three.
Read next