Sutra

The physics test: what video models still break, and how to check in ten seconds

Eleven things current video models still get wrong, from liquids to the frame an object is released on. What to look at, what the failure looks like, and whether it can be fixed in the edit or the shot has to be made again.

What is in here
  1. Why does AI video break physics?
  2. The eleven tests, grouped by what is actually being violated
  3. Why is a wrong event caught faster than a wrong shadow?
  4. How do I check all of this in ten seconds?
  5. What does our own physics check actually ask?
  6. What this battery cannot do
The short answer

Video models are good at appearance and bad at consequence. The eleven places they still break are all relational: liquid volume, cloth weight, grip, contact shadow, occlusion order, mass and inertia, reflections, hair, granular scatter, packaging text, and the exact frame an object is released on. You can check every one of them on a single paused frame in about ten seconds. The verdict that matters is whether the wrongness can be cut around, or whether the shot has to be generated again.

What you get out of this
  1. Eleven named tests, each with the thing to look at and the failure that gives it away
  2. A verdict for every one: fixable in the edit, or the shot gets made again
  3. Why your eye catches a wrong event faster than it notices a wrong shadow
  4. The question our own physics check asks, which is not the question you would guess

Six places physics is decided, in one frame we shipped

Tap the numbers
A gooseneck kettle pouring a thin stream of water into a dark ceramic pour-over cone full of coffee bloom, sitting on a glass carafe half full of coffee, on a wooden table
One of our own frames, from a cafe demo film. Nothing here is a failure. These are the six places a failure would have been visible, which is why they are worth stopping on.

01Why does AI video break physics?

Because a video model predicts what the next frame usually looks like, not what the world would do next. It has no volume, no mass and no contact solver. What it gets right about physics is an echo of footage where physics was already obeyed, and echoes fail where your product is unusual.

The academic version is real and active: benchmarks like VideoPhy score models on physical commonsense, and the search results for physics in generated video are almost entirely papers. What none of them tell you is which failure costs you a shot on Thursday. The breaks cluster where two things interact, because that is the only place a model has to satisfy two constraints at once, and it will satisfy whichever one was commoner in its training. Hands are the worst case, which is why they get a piece of their own.

02The eleven tests, grouped by what is actually being violated

Pause the shot on a frame where two things are touching. Ten of the eleven live there. Sorting them by what they look like gives you a list; sorting them by which physical law is being broken gives you a triage, because the fix for a broken constraint is nothing like the fix for a broken material.

The battery: five families, eleven tests

Switch between them
Liquids and granular materials

Watch the leading edge of a pour, not the splash. A real stream narrows as it accelerates and breaks into droplets at a distance you can estimate from the height it fell. Generated streams tend to hold a constant width the whole way down, because width is a texture and narrowing is a consequence of something the model is not computing.

Granular material is the same test at a smaller scale, and it is much less forgiving. Salt, ground coffee, powder and sand separate into individual paths the moment they leave a hand. When a model gets it wrong the pinch falls as one soft sheet, more like fabric than like grains, and lands without scattering.

  • Volume that is not conserved: the glass fills faster than the stream delivers
  • A stream that never breaks into droplets anywhere along its length
  • A surface that stays flat while something is entering it
  • Grains that land as a single mass, with no bounce and no spread
Look at
the leading edge of the pour, and the surface where it lands
Verdict
Re-roll. There is no post fix for volume.
Every one of these came out of a frame we either shipped, killed or re-rolled. The verdict line is the part worth arguing with, because it is the one that spends your money.

Eleven tests, and what each failure costs you

The verdict column
Eleven tests, and what each failure costs you
DimensionWhat the failure looks likeFixable in the edit
LiquidsVolume appears or vanishes mid-pourNo. Re-roll
Cloth and fabric fallThe fabric moves at the wrong weightNo. Re-roll
Hands on objectsFingers grip nothing, or pass throughNo. Re-roll, or cut away
Contact shadowsDetached, or offset by a fixed amountPartly. Often, if the object is still
Occlusion orderThe near thing goes behind the far thingNo. Re-roll
Mass and inertiaA heavy object moves like an empty oneNo. Re-roll
Reflections and specularsThe reflection disagrees with its objectPartly. Sometimes, with a crop
HairStrands merge, split or gain volumePartly. Only by shortening the shot
Granular materialsGrains fall and land as one sheetNo. Re-roll
Legible packaging textYour own label reads as nonsenseYes. Yes, with a real insert
Gravity at releaseThe object hangs a beat, then dropsNo. Re-roll, or cut earlier
The right column is the only one that changes a schedule. Ours, from the way we actually triage a failed shot: appearance is repairable and behavior is not, and a shot that fails on behavior is re-rolled rather than nursed.

The physics check does not ask whether an event happened. It asks whether it was designed, or cheap.

The verdict our own checker returns, in one word

03Why is a wrong event caught faster than a wrong shadow?

Because an event is predicted and a shadow is only observed. Battaglia, Hamrick and Tenenbaum argued in PNAS that people understand physical scenes by running an approximate simulation forward, which means a viewer has already computed what should happen next before the next frame arrives. A violated prediction registers as a small failure, fast and pre-verbal. Nobody predicts a shadow. You have to go looking for it.

A hand held above a sage green ceramic plate with fingers pinched, releasing shredded garnish over grilled fish, shaved onion, red sauce and a wedge of orange, in a dark dining room with a candle out of focus behind
The hardest half-second in this frame is the one that has not happened yet. Everything the pinch does after the fingers open is a physics problem, and the release frame is where it is decided.

That asymmetry is worth real money, because it inverts the order most people work in. Grade, grain and shadow craft are what a maker fusses over, and they are the slow half of the viewer's read. The event is the fast half. We learned this the expensive way, on a shot of a photographed loaf pulling apart that cleared every automated gate we owned and was killed on sight for behaving in a way bread does not. The whole incident, and the fifteen other tells that came out of it, are in the frame-by-frame atlas of why AI ads look fake.

The film the hotspot frame came from. Every event in it is real footage, because a pour is the one thing on this page nobody should be asking a model to invent.
The rule this battery came out ofA still or a cutout may move as a card: push, settle, drift, a light patch traveling across it. It may never perform a physical event it would have to invent geometry for. If the treatment is openly designed and reads as illustration, the event is legal, because nothing is pretending to be a photograph. The case for choosing that route deliberately is in when a declared style beats a bad photoreal.

04How do I check all of this in ten seconds?

Scrub to a frame where two things are touching, and hold there. Contact answers six of the eleven tests at once, because grip, shadow, occlusion, weight, reflection and material all resolve in the same few hundred pixels. Then step forward one frame at a time through the nearest release. That is the whole pass, and it is faster to run than to describe.

The ten-second pass, on a cut you already have

Tick as you go - it remembers
0%
Six checks. Run them cold, on a phone, at the size the ad will actually play. If you cannot honestly tick a box, that is the note, and you do not need a second opinion to act on it.

What to do with a failed test

The fork is always the same, and it is not a taste question. Appearance is repairable: grade, grain, crop, speed and sound can all be changed after the fact. Behavior is baked into the pixels. Contact, deformation, occlusion order and depth need the shot made again, and a shot that needs making again is re-rolled rather than nursed toward acceptability, which is a budget rule before it is a craft rule.

What happens after a test fails

The fork
01Name the testwhich of the eleven actuallyfailed, in one sentence02Appearance or behaviorgrade, grain and crop areappearance. Contact is notTHE FORK03Cut around itshorten the shot, crop thecorner, cut before therelease04Re-rollgenerate the shot again.Never repair a bad takeNEEDS A YES
01Name the testwhich of the eleven actually failed, in one sentence
02Appearance or behaviorgrade, grain and crop are appearance. Contact is notthe fork
03Cut around itshorten the shot, crop the corner, cut before the release
04Re-rollgenerate the shot again. Never repair a bad takeneeds a yes
Two of these four steps cost nothing. The fourth is the only one that spends money, which is why it needs a yes from whoever is paying rather than from whoever is building.

05What does our own physics check actually ask?

Not whether an event occurred. That question has an answer and the answer is useless. The pass asks whether the event reads as designed or as cheap, and it is answered by somebody who never touched the build and sees only the delivered file. A maker cannot run it on their own work, because a maker knows what the shot was supposed to be doing and will see it whether it is there or not.

A physics verdict, laid out as a transcript

Watch it run
This is a written verdict from a checker who sees only the delivered bundle, set out here as a transcript. It is not a command you can run, and we would rather say so: our engine has no eye. Every physics defect we have ever caught was caught by rendering frames and looking at them.

06What this battery cannot do

It cannot tell you whether an ad works, because nothing in our reference set carries a hook rate or a return on spend to check a shot against. Fewer than half the films that clear every automated gate we run survive a human watching them afterwards, which is the honest shape of the gap between passing a check and being good.

Four numbers behind the battery

Ours, counted off delivered files
11physics tests we run on a single paused frame before a generated shot is acceptedSutra Haus
43%of the films that clear every automated gate we run are still killed by a human watching themSutra Haus production log
0.83smedian average shot length across the 48 reference films we measured. Pace is what stops a viewer examining a generated frameSutra Haus corpus study
2.5sour default ceiling on a single picture beat, raised only when a shot has an internal arc that earns the timeSutra Haus
All four are ours. The second one is the uncomfortable one, and it is the reason we treat every automated check as a floor rather than as a verdict.

It also cannot be automated, and we have tried. Every substitute for judgment we have built failed toward a confident wrong answer: four detectors, five limb measurers, a cut counter wrong by nine times in one direction, and a frame-delta gate that ranked our accepted films below our rejected ones and was withdrawn. Use the battery as a floor and keep the ceiling in your own eyes. There is a wider version of this pass, covering type and sound as well, in the pre-launch QA checklist for generated footage.

Questions people actually ask

Open what you need
Why does AI video break physics?

A video model predicts the next frame, not the next state of the world. It carries no representation of mass, volume or contact, so anything it gets right is inherited from footage where physics was already being obeyed. That inheritance holds up for common objects doing common things, and fails at the edges: your product, your fabric, your grip, your pour.

How do I tell if a video is AI generated?

Look at behavior rather than texture. Pause on a frame where two things touch and check that both of them change. Then step through the nearest moment an object is released and confirm it starts falling on the next frame.

Texture is the first thing people mention and the last thing that decides it. We have shipped clean-textured shots that died on a wrong event, and rough ones nobody questioned.

Is there an AI video physics test I can actually run?

Yes, and it takes about ten seconds per shot. Pause on contact, follow one lock of hair or one fold of cloth for ten frames, find the release and step through it, confirm the near object covers the far one for the whole shot, and read your own packaging text at full size. The checklist on this page is that pass, in the order we run it.

Can I fix bad AI physics in post?

Appearance is fixable, behavior is not. Grade, grain, crop, speed and sound can all be changed afterwards. Contact, deformation, occlusion order and depth are baked into the pixels. In practice that means shadows and reflections are often rescuable, hair is rescuable by shortening the shot, packaging text is rescuable with a real photograph dropped in, and everything else on the list gets generated again.

Which of these will the next model fix?

Some of them, and the ones you would expect: streams, foam, bubbles and hair have all improved fast, because they are common in training footage. Grip on an unfamiliar object and text on your specific label are structurally harder, because no amount of general footage contains your product. Those two are worth designing around rather than waiting on.

Does slowing the footage down help?

It usually makes things worse. Slowing a clip stretches the frames where the physics is wrong and gives a viewer longer to run the prediction that catches it. Fast cutting protects generated footage for exactly the opposite reason. If a shot only survives at speed, it is a short shot, not a slow one.

None of these eleven get solved by waiting for a better release, because none of them is about fidelity. Every one is a question about something the model was never told: what your product weighs, how your fabric hangs, how many grams are in that pinch, what your label says. Tell it, shoot it, or cut the shot. Those are the only three moves, and the third one is free.

Where the numbers came from

  1. PNAS. Simulation as an engine of physical scene understanding (Battaglia, Hamrick and Tenenbaum, 2013) - the argument that people predict physical scenes by running an approximate simulation, which is why a wrong event is caught before it can be described
  2. arXiv. VideoPhy: Evaluating Physical Commonsense for Video Generation - one of the academic benchmarks for this problem. The search results for physics in generated video are almost entirely papers like this one, which is why we wrote the working version
  3. IEEE Spectrum. The Uncanny Valley (Masahiro Mori, translated by MacDorman and Kageki) - why a near-miss costs more than an obvious miss

Every figure above links to the place it was published. Numbers marked as ours are measured inside this studio and we say so where they appear. We do not print a statistic we cannot point at.

Badal Kariwal

Runs Sutra Haus, a one-person ad studio that has shipped over a thousand finished creatives - film and stills - for DTC brands and hotels. Writes here about what the work actually taught him, including the parts that failed. The person who reads your brief is the person who builds the work. Send him something to make.

Run the battery on something of ours

Send a link. Get one finished ad back.

One finished cut built from your own product, inside three days, free and yours to run whether or not we ever work together. Pause it on contact and run the eleven tests. If it fails one, you will know before we do.

Replies within a day. Ad within three.
Read next