Sutra

Legibility floors: the smallest caption that still works on a phone

There is no correct font size for a subtitle, and we proved it against our own work: fourteen of fifteen delivered cards failed a re-measure. Two floors decide it instead, and you can check both on a file you already shipped.

What is in here
  1. Is there a standard font size for subtitles?
  2. Floor one: contrast under the worst pixel, not the average
  3. How long does a caption need to stay on screen?
  4. Weight beats size, and it is a resolution fact
  5. A plate is not a failure of nerve
  6. How do I check my own captions?
The short answer

There is no standard font size, and any number somebody gives you is measured on the wrong thing. Two floors decide it. Contrast: at least 4.5 to 1 against the worst local background under each run of letters, not the average, checked at delivery size and again at 360 pixels. Dwell: every line holds at full opacity for at least 1.2 seconds, or 0.35 seconds per word, whichever is longer. Fades do not count.

What you get out of this
  1. Why a percentage-of-frame-height rule is not enough, proved by two cards at the identical size with opposite outcomes
  2. How to measure contrast the way the eye does it: locally, under each glyph run, at embed width
  3. The dwell arithmetic, and a tool below where you can type your own line and watch it fail
  4. Why a size bump cannot fix thin type, with the sub-pixel number that explains it
  5. What we found when we re-measured fifteen cards we had already delivered

What happened when we checked our own delivered files

One re-measure of work we had already shipped
14/15delivered cards that failed the re-measureSutra Haus legibility audit
60/78individual runs of text that failed inside those cardsSutra Haus legibility audit
10runs that lost their stroke completely at 360 pixels. The letters had physically dissolvedSutra Haus legibility audit
1card that passed, and the only one that did not put type over a photographSutra Haus legibility audit
We rebuilt fifteen delivered static creatives, confirmed the rebuilds were byte-identical to what had shipped, then measured them again with a gate that reads the worst local background under each run of letters instead of an average, in perceptual terms rather than raw pixel values, and again at 360 pixels. These four numbers are the result.

01Is there a standard font size for subtitles?

No, and we ran one for months before our own numbers killed it. The rule was a readability floor expressed as a percentage of frame height, which sounds rigorous and is measuring the wrong property. One card passed at 0.78 percent of frame height and was legible without effort: black ink on pale green, 150 pixels tall. Its sibling was unreadable at the identical 150 pixels and the identical 0.78 percent, because it was dark blue on pale blue.

Same diameter, same percentage, opposite outcomes. Legibility is ink contrast and source glyph resolution. It is not delivered size.

The note that retired our own readability floor

Three ways people measure caption size, and what each one misses

What each one can and cannot tell you
Three ways people measure caption size, and what each one misses
DimensionWhat it catchesWhat walks straight past it
Point size in your editorAlmost nothing useful.No. Cap height varies by face. A display serif can sit at about 0.31 of its point size, less than half a grotesque. We set chapter words at 62pt and measured 2.40 percent cap height, a 15-pixel smudge in a feed tile.
Percentage of frame heightGenuinely useful as a floor. Ours is 5 percent cap height, measured off the ink.No. Every contrast failure. It passed our blue-on-blue card at exactly the same number as the card next to it that worked.
Average contrast per stringObvious black-on-black mistakes.No. Local failure. One of our headlines averaged 0.356 and measured 1.17 to 1 against the actual worst ground under the letters, because the word sat across a doorway edge and a lit white blouse.
Worst local ground, per glyph run, at two widthsYes. What we run nowType over moving footage, which we sample one frame per text moment. A caption can pass on the sampled frame and fail two seconds later. We say so because it is still open.
We now measure two things on the delivered file and infer nothing from a point size: cap height as a fraction of frame height, read off the plate's own ink, and perceptual contrast against the ring of pixels around each run of letters.

The two thresholds themselves are not ours. They are the W3C's: 4.5 to 1 for normal text and 3 to 1 for large text. What is ours is where we apply them. Full 4.5 at delivery width, relaxed to 3 at 360 pixels, because most of the loss at embed size is a resampling artifact rather than a design fault, and holding the tighter number there would fail things that are genuinely fine.

02Floor one: contrast under the worst pixel, not the average

A mean of 36 and 239 describes no pixel anybody is looking at. That sentence is the whole section. If a word crosses from a dark doorway onto a lit shoulder, the average says the contrast is comfortable and the second half of the word is invisible. Measure under each run of letters, take the worst ground you find there, and judge against that.

One word, four different backgrounds

Tap the numbers
A woman in a black strap top seen in profile, warm lamplight on her shoulder, a circular patch on her upper arm, a dark bookcase behind her, and the word PRESS set large in heavy white type across the lower frame
This frame passes, and it passes for reasons that are decisions rather than luck: one word, heavy weight, and a set of grounds that are all dark enough at the point the ink actually sits.

When we re-measured fifteen delivered cards, two patterns accounted for nearly all the failures. The first was pale display type over a photograph with a soft gradient scrim under it: a scrim is a ramp, so it is weakest exactly where the type sits. The second was small letterspaced capitals at 24 to 30 pixels, which look considered on a desktop and stop existing in a tile.

The check we run before anything ships

Watch it run
A simplified transcript of our own legibility pass. The two widths matter: a run can clear the gate at delivery size and lose its stroke entirely at embed size, and the second number is the one your buyer sees.
The rule that keeps this honestA lint is a floor, not a target. We once measured a scrim at 0.299 against a floor of 0.300, one thousandth under, on a delivered file. The scrim got longer. Nobody touched the floor. Tuning a threshold until the answer looks right is fitting to noise, and we have the receipts on where that ends: we once tuned a soundtrack until it satisfied a frequency check and shipped a bed the client described as physically painful to listen to.

03How long does a caption need to stay on screen?

At least 1.2 seconds, or 0.35 seconds per word, whichever is longer, measured at full opacity. A six-word line therefore needs 2.1 seconds. Fades are excluded from the count, because type at partial alpha is not readable, and excluding them is the difference between a rule and a rule that works.

Will your caption survive a phone?

Type your own line into it
A hand seasoning a plated dish from above, a pale green ceramic plate in the foreground and a blown-out bright window behind, from one of our hotel films

Words-
Dwell floor-
Your hold-
Contrast-
The dwell arithmetic here is exactly the gate we run: max(1.2s, 0.35s per word). The contrast readout is a simplification of it. Ours measures the worst local ground under every run of letters rather than one figure for the frame, which is the difference that failed fourteen of our own cards.

Two word rates sit underneath that floor and they are three to four times apart. A word that replaces what was on screen needs roughly 0.75 seconds. A word that joins a string already there needs only 0.2 to 0.4, because by the fifth cut the reader has had a second and a half with the sentence and has one new word to pick up. Confusing the two is how films end up moving faster than their own content.

We watched exactly that happen. One film put eleven screens of paragraph copy across 279 frames, a mean of 0.85 seconds each, and not one of them was readable. The type was beautiful. The pacing was confident. The film moved faster than its own content, which is a sentence I now use as a diagnosis rather than a description.

04Weight beats size, and it is a resolution fact

It sounds like a taste argument and it is arithmetic. An ultralight face at 36 pixels integrates to a stem 0.71 pixels wide. That is narrower than a pixel, so no stroke ever fully covers a pixel column, and only 6 to 7 percent of the letter's ink reaches full opacity. The same face at regular weight has a 3.70 pixel stem and 78 percent of its ink at full.

What weight does that size cannot

The same face, the same size
What weight does that size cannot
DimensionUltralight at 36pxRegular at 36px
Stem width after renderingNo. 0.71 pxYes. 3.70 px
Share of ink at full opacityNo. 6 to 7 percentYes. 78 percent
What a size bump does to itBigger, and equally faint. The stem scales with the size and the sub-pixel problem travels with it.Nothing needed.
What it does at 360 pxNo. Dissolves. We lost the crossbar of every H on one card and the word read as PATCIIESYes. Holds
Both rows are the same typeface at the same 36 pixels. The only variable is weight. This is why we treat a size bump as the wrong reflex: it makes a faint letter bigger and just as faint.

A related trap, because it is the same mistake wearing a nicer suit: one weight of one face is about ratio, not count. A 134-pixel ultralight headline sitting above a 42-pixel light sub-line reads as two different typefaces to anybody who is not the person who set it.

05A plate is not a failure of nerve

Designers resist putting a solid shape behind type because it looks like a compromise. It is not a compromise, it is the only way to make the ground under the letters a decision instead of an accident. Where the plate is allowed to sit is a separate question, and the safe zones are constraints rather than guidelines. The single card that passed our re-measure was the only one of fifteen that did not put type over a photograph. That is not a coincidence and it is not a small finding: the failure was the house style, not one bad card.

Type dies on a cut. Never on a fade.

We confirmed this to the frame, three times inside a single reference film: the title vanishes at frame 53, the ingredient list at frame 290, a headline at frame 415, and all three of those frames are picture cuts. Nothing fades out. Type appears inside a shot and is killed by the edit, which is also why fade time cannot be counted toward dwell.

Five things about captions we have argued about and lost

Flip them
The first three are ours. The last two are things clients have said to us that turned out to be right, which is worth admitting in a piece about being wrong.

The words a legibility argument needs

Search the vocabulary
6 terms
Cap heightMeasurement
The height of a capital letter as it renders, measured off the ink rather than inferred from a point size. Our floor is 5 percent of frame height. Point size will not tell you this, because cap height varies enormously between faces.
Glyph runMeasurement
One continuous stretch of letters that shares a background. Contrast is measured per run and against the worst ground inside it, because a line crossing two grounds is really two problems wearing one typeface.
Worst local groundContrast
The single most punishing patch of picture sitting behind any part of a run of letters. It is the number the eye actually experiences, and it is the number an average is designed to hide.
Dwell floorTiming
The minimum time a line holds at full opacity, computed from its own word count: 1.2 seconds, or 0.35 seconds per word, whichever is greater. Fade time never counts, because type at partial alpha is not readable.
PlateCraft
A solid or near-solid shape placed under type so the ground behind the letters is chosen rather than inherited. Distinct from a scrim, which is a gradient, and therefore weakest exactly where the type sits.
Stroke coverageRendering
How much of a pixel column a letter stem actually fills after rendering. Below one pixel, no part of the stroke ever reaches full opacity, and the letter reads as a smear. This is a weight problem and a size increase cannot fix it.
Six terms that make this conversation possible. Most caption arguments are really two people using the word size to mean different things.

06How do I check my own captions?

Export the finished file, not the timeline. Scale it to 360 pixels wide, which is roughly the real width of a feed tile, and look at it there. Do it in the same pass as the rest of the second-by-second build check, because type timing and cut timing are the same timing. Everything below is doable in ten minutes without buying anything, and the first two checks will find most of what is wrong.

The ten-minute caption audit

Tick as you go - it remembers
0%
Run it on the delivered file, at 360 pixels, on a phone, in daylight. Not on your timeline at 100 percent, which is the condition under which every one of the failures in this piece looked fine.

Questions people actually ask

Open what you need
What is the minimum text size for an Instagram Reels ad?

There is no single answer in pixels, because pixel size does not survive the resize to a feed tile and different typefaces produce wildly different cap heights at the same point size. Our own floor is cap height at 5 percent of frame height or more, measured off the rendered ink rather than the point size. On a 1920-pixel-tall frame that is about 96 pixels of cap height, which is far larger than most people expect.

Is there a standard font size for subtitles?

Not one that transfers. Broadcast subtitling has had reading-rate and layout rules for decades, and the BBC publishes its own, but they were written for a fixed screen at a fixed distance. A social ad is a 360-pixel tile in daylight in one hand.

What does transfer is the pair of properties underneath every one of those standards: enough contrast against whatever is actually behind the letters, and enough time on screen to read them. Measure those two and the size question answers itself.

How do I make captions readable over a busy video?

Put a real shape behind them. A solid or near-solid plate makes the ground under the letters a decision rather than an accident, and a gradient scrim does not, because a gradient is weakest exactly where the type sits. If a plate is genuinely wrong for the piece, move the type to the most stable part of the frame and check it against the worst pixel it crosses, not the average.

How long should a subtitle stay on screen?

At least 1.2 seconds, or 0.35 seconds per word, whichever is longer, at full opacity. A six-word line needs 2.1 seconds. The exception runs the other way: a word joining a sentence already on screen can appear for 0.2 to 0.4 seconds, because the reader has already read the rest of it and is acquiring one new word.

Do automated caption checkers work?

As a prompt to look, yes. As a verdict, no. In one census of 64 files, band detection returned 24 false positives, several of them a perfect 100 percent score, caused by bright walls, white clothing, high-key grades and ceiling lights. Every automated substitute for looking that we have built has failed toward a confident wrong answer rather than an obviously broken one, which is the dangerous direction.

Does any of this matter if people watch with sound on?

Yes, for a different reason. Captions that fail are not neutral, they are decorative text nobody can read, and that is a specific and visible defect. It also happens to be the cheapest thing in the film to get right, which is why it reads so badly when it is wrong. Sound is a separate argument, and briefing it for a muted feed is its own discipline.

Everything in this piece is a floor. Clearing all of it produces a caption that can be read, which is not the same as a caption worth reading, and no gate we own has an opinion about the second thing. But a line nobody can read has no chance at all, and it is the cheapest failure in the entire film to prevent. Fix the type, never the threshold.

Where the numbers came from

  1. W3C. Understanding Success Criterion 1.4.3: Contrast (Minimum) - the 4.5:1 normal-text and 3:1 large-text thresholds both floors are built from
  2. BBC. Subtitle Guidelines - broadcast subtitling has had reading-rate rules for decades; the principle predates social video by a long way

Every figure above links to the place it was published. Numbers marked as ours are measured inside this studio and we say so where they appear. We do not print a statistic we cannot point at.

Badal Kariwal

Runs Sutra Haus, a one-person ad studio that has shipped over a thousand finished creatives - film and stills - for DTC brands and hotels. Writes here about what the work actually taught him, including the parts that failed. The person who reads your brief is the person who builds the work. Send him something to make.

Have us run this on something you already shipped

Send a link. Get one finished cut back.

One 9:16 film built from your own product, inside three days, free and yours whether or not we ever work together. Scale it to 360 pixels and check every line in it against the two floors in this piece. That is exactly the check it had to clear before you saw it.

Replies within a day. Ad within three.
Read next