The physics test: what video models still break, and how to check in ten seconds
Eleven things current video models still get wrong, from liquids to the frame an object is released on. What to look at, what the failure looks like, and whether it can be fixed in the edit or the shot has to be made again.
What is in here
Video models are good at appearance and bad at consequence. The eleven places they still break are all relational: liquid volume, cloth weight, grip, contact shadow, occlusion order, mass and inertia, reflections, hair, granular scatter, packaging text, and the exact frame an object is released on. You can check every one of them on a single paused frame in about ten seconds. The verdict that matters is whether the wrongness can be cut around, or whether the shot has to be generated again.
- Eleven named tests, each with the thing to look at and the failure that gives it away
- A verdict for every one: fixable in the edit, or the shot gets made again
- Why your eye catches a wrong event faster than it notices a wrong shadow
- The question our own physics check asks, which is not the question you would guess
Six places physics is decided, in one frame we shipped
Tap the numbers
01Why does AI video break physics?
Because a video model predicts what the next frame usually looks like, not what the world would do next. It has no volume, no mass and no contact solver. What it gets right about physics is an echo of footage where physics was already obeyed, and echoes fail where your product is unusual.
The academic version is real and active: benchmarks like VideoPhy score models on physical commonsense, and the search results for physics in generated video are almost entirely papers. What none of them tell you is which failure costs you a shot on Thursday. The breaks cluster where two things interact, because that is the only place a model has to satisfy two constraints at once, and it will satisfy whichever one was commoner in its training. Hands are the worst case, which is why they get a piece of their own.
02The eleven tests, grouped by what is actually being violated
Pause the shot on a frame where two things are touching. Ten of the eleven live there. Sorting them by what they look like gives you a list; sorting them by which physical law is being broken gives you a triage, because the fix for a broken constraint is nothing like the fix for a broken material.
The battery: five families, eleven tests
Switch between themLiquids and granular materials
Watch the leading edge of a pour, not the splash. A real stream narrows as it accelerates and breaks into droplets at a distance you can estimate from the height it fell. Generated streams tend to hold a constant width the whole way down, because width is a texture and narrowing is a consequence of something the model is not computing.
Granular material is the same test at a smaller scale, and it is much less forgiving. Salt, ground coffee, powder and sand separate into individual paths the moment they leave a hand. When a model gets it wrong the pinch falls as one soft sheet, more like fabric than like grains, and lands without scattering.
- Volume that is not conserved: the glass fills faster than the stream delivers
- A stream that never breaks into droplets anywhere along its length
- A surface that stays flat while something is entering it
- Grains that land as a single mass, with no bounce and no spread
- Look at
- the leading edge of the pour, and the surface where it lands
- Verdict
- Re-roll. There is no post fix for volume.
Cloth, hair, and the frame of release
Fabric weight is legible in exactly one thing: how far a wave travels along the cloth before it dies. Heavy fabric moves in a few large folds and settles fast. Light fabric ripples across its whole area and keeps going after the hand has stopped. A generated garment often drapes correctly in a still and then moves at the wrong weight the moment anything touches it.
Hair breaks the most quietly of anything on this list. Follow one lock across ten frames. Strands should stay a countable group that keeps its own volume; when they merge, split, or grow between frames, the shot has been improvising and you have found the seam.
Gravity is easiest to catch at the exact frame of release. Something leaves a hand and should begin falling in the very next frame, slowly, then faster. A held beat before the fall is the most common single physics defect we see, and people almost always describe it as slow motion when what is actually missing is acceleration.
- A wave that travels the whole length of a garment too heavy to allow it
- Hair that changes strand count between frames
- An object that hangs for two or three frames after the fingers open
- A fall at constant speed rather than an accelerating one
- Look at
- one lock of hair, one fold of cloth, and the single frame of release
- Verdict
- Re-roll, unless the shot can be cut before the release.
Hands on objects, contact shadows, occlusion order
Contact is a two-way event. A hand takes a cup and the fingers flatten a little, the cup does not, and the skin changes color where the pressure is. Generated contact is usually one-way: a hand deforming around nothing, or an object held by fingers that never closed.
The contact shadow is the fastest read in the frame. There should be two of them, not one: a tight dark shadow where the object meets the surface, and a wider ambient one further out. The offset between object and shadow scales with the object's own width, which is exactly why one shadow recipe copied onto a bigger object immediately reads as pasted on.
Occlusion order is the cheapest test here and the one almost nobody runs. Take the two things nearest the camera and confirm the nearer one covers the further one for the entire shot. Generated footage swaps them across a cut more often than you would expect, and once you have seen it you cannot stop seeing it.
- Fingers that curve around an object without ever touching it
- One soft blob of shadow under the whole object
- A shadow that keeps the same offset as the object gets larger
- A near object briefly rendered behind a far one
- Look at
- the exact pixels where two things meet
- Verdict
- Shadows are often fixable. Grip and occlusion are not.
Mass and inertia
Weight is never in the object. It is in the body carrying it, and in what happens three frames after the motion stops. Look at the wrist, the shoulder and the opposite hip, not at the thing being lifted. If the person would look identical holding nothing, the weight is a costume.
Inertia is the other half. A full bottle set down hard keeps its liquid moving after the glass has stopped. A heavy door swings past its resting point and comes back. Generated motion tends to arrive at rest on precisely the frame the movement ends, which reads as an animation preset rather than an event. We ease every move we build with smoothstep and nothing else, with zero overshoot and zero bounce, for the same reason: a photographed object has no mass to borrow.
- A carried object that does not tilt the person carrying it
- Liquid that stops on the same frame as its container
- A settle with a spring or a bounce on it, which no photographed object has
- Look at
- the body, and the three frames after the motion stops
- Verdict
- Re-roll. Weight cannot be added in the edit.
Reflections, specular highlights, packaging text
A reflection has to agree with its object about position, angle and softness, and it has to fade at the rate the surface roughness implies. Specular highlights have a stricter job than that: they must move when the camera moves and stay put when only the object moves, because a highlight is a fact about the light rather than about the object.
Packaging text is where generated footage gives itself away to your customer rather than to you. A model has never read your label, so it writes something label-shaped. Anyone who owns the product can see it in the first second. This is the one item on the list with a clean fix, and the fix is a photograph of the real pack, dropped in as a card and moved like a card. The related trap is borrowed footage carrying somebody else's product name; we once had to cap every window on a stock clip because a competitor's label became legible above a certain magnification, and the shot we wanted was simply lost.
- A highlight that stays in the same place on the object through a camera move
- A reflection that is a copy of the object rather than a view of it
- Label text legible enough to read, and wrong
- Somebody else's brand name readable in footage you borrowed
- Look at
- the brightest pixel, and any text taller than about twelve pixels
- Verdict
- Text is fixable with a real insert. Reflections usually are not.
Eleven tests, and what each failure costs you
The verdict column| Dimension | What the failure looks like | Fixable in the edit |
|---|---|---|
| Liquids | Volume appears or vanishes mid-pour | No. Re-roll |
| Cloth and fabric fall | The fabric moves at the wrong weight | No. Re-roll |
| Hands on objects | Fingers grip nothing, or pass through | No. Re-roll, or cut away |
| Contact shadows | Detached, or offset by a fixed amount | Partly. Often, if the object is still |
| Occlusion order | The near thing goes behind the far thing | No. Re-roll |
| Mass and inertia | A heavy object moves like an empty one | No. Re-roll |
| Reflections and speculars | The reflection disagrees with its object | Partly. Sometimes, with a crop |
| Hair | Strands merge, split or gain volume | Partly. Only by shortening the shot |
| Granular materials | Grains fall and land as one sheet | No. Re-roll |
| Legible packaging text | Your own label reads as nonsense | Yes. Yes, with a real insert |
| Gravity at release | The object hangs a beat, then drops | No. Re-roll, or cut earlier |
The physics check does not ask whether an event happened. It asks whether it was designed, or cheap.
The verdict our own checker returns, in one word
03Why is a wrong event caught faster than a wrong shadow?
Because an event is predicted and a shadow is only observed. Battaglia, Hamrick and Tenenbaum argued in PNAS that people understand physical scenes by running an approximate simulation forward, which means a viewer has already computed what should happen next before the next frame arrives. A violated prediction registers as a small failure, fast and pre-verbal. Nobody predicts a shadow. You have to go looking for it.

That asymmetry is worth real money, because it inverts the order most people work in. Grade, grain and shadow craft are what a maker fusses over, and they are the slow half of the viewer's read. The event is the fast half. We learned this the expensive way, on a shot of a photographed loaf pulling apart that cleared every automated gate we owned and was killed on sight for behaving in a way bread does not. The whole incident, and the fifteen other tells that came out of it, are in the frame-by-frame atlas of why AI ads look fake.
04How do I check all of this in ten seconds?
Scrub to a frame where two things are touching, and hold there. Contact answers six of the eleven tests at once, because grip, shadow, occlusion, weight, reflection and material all resolve in the same few hundred pixels. Then step forward one frame at a time through the nearest release. That is the whole pass, and it is faster to run than to describe.
The ten-second pass, on a cut you already have
Tick as you go - it remembersWhat to do with a failed test
The fork is always the same, and it is not a taste question. Appearance is repairable: grade, grain, crop, speed and sound can all be changed after the fact. Behavior is baked into the pixels. Contact, deformation, occlusion order and depth need the shot made again, and a shot that needs making again is re-rolled rather than nursed toward acceptability, which is a budget rule before it is a craft rule.
What happens after a test fails
The fork05What does our own physics check actually ask?
Not whether an event occurred. That question has an answer and the answer is useless. The pass asks whether the event reads as designed or as cheap, and it is answered by somebody who never touched the build and sees only the delivered file. A maker cannot run it on their own work, because a maker knows what the shot was supposed to be doing and will see it whether it is there or not.
A physics verdict, laid out as a transcript
Watch it run06What this battery cannot do
It cannot tell you whether an ad works, because nothing in our reference set carries a hook rate or a return on spend to check a shot against. Fewer than half the films that clear every automated gate we run survive a human watching them afterwards, which is the honest shape of the gap between passing a check and being good.
Four numbers behind the battery
Ours, counted off delivered filesIt also cannot be automated, and we have tried. Every substitute for judgment we have built failed toward a confident wrong answer: four detectors, five limb measurers, a cut counter wrong by nine times in one direction, and a frame-delta gate that ranked our accepted films below our rejected ones and was withdrawn. Use the battery as a floor and keep the ceiling in your own eyes. There is a wider version of this pass, covering type and sound as well, in the pre-launch QA checklist for generated footage.
Questions people actually ask
Open what you needWhy does AI video break physics?
A video model predicts the next frame, not the next state of the world. It carries no representation of mass, volume or contact, so anything it gets right is inherited from footage where physics was already being obeyed. That inheritance holds up for common objects doing common things, and fails at the edges: your product, your fabric, your grip, your pour.
How do I tell if a video is AI generated?
Look at behavior rather than texture. Pause on a frame where two things touch and check that both of them change. Then step through the nearest moment an object is released and confirm it starts falling on the next frame.
Texture is the first thing people mention and the last thing that decides it. We have shipped clean-textured shots that died on a wrong event, and rough ones nobody questioned.
Is there an AI video physics test I can actually run?
Yes, and it takes about ten seconds per shot. Pause on contact, follow one lock of hair or one fold of cloth for ten frames, find the release and step through it, confirm the near object covers the far one for the whole shot, and read your own packaging text at full size. The checklist on this page is that pass, in the order we run it.
Can I fix bad AI physics in post?
Appearance is fixable, behavior is not. Grade, grain, crop, speed and sound can all be changed afterwards. Contact, deformation, occlusion order and depth are baked into the pixels. In practice that means shadows and reflections are often rescuable, hair is rescuable by shortening the shot, packaging text is rescuable with a real photograph dropped in, and everything else on the list gets generated again.
Which of these will the next model fix?
Some of them, and the ones you would expect: streams, foam, bubbles and hair have all improved fast, because they are common in training footage. Grip on an unfamiliar object and text on your specific label are structurally harder, because no amount of general footage contains your product. Those two are worth designing around rather than waiting on.
Does slowing the footage down help?
It usually makes things worse. Slowing a clip stretches the frames where the physics is wrong and gives a viewer longer to run the prediction that catches it. Fast cutting protects generated footage for exactly the opposite reason. If a shot only survives at speed, it is a short shot, not a slow one.
None of these eleven get solved by waiting for a better release, because none of them is about fidelity. Every one is a question about something the model was never told: what your product weighs, how your fabric hangs, how many grams are in that pinch, what your label says. Tell it, shoot it, or cut the shot. Those are the only three moves, and the third one is free.
Where the numbers came from
- PNAS. Simulation as an engine of physical scene understanding (Battaglia, Hamrick and Tenenbaum, 2013) - the argument that people predict physical scenes by running an approximate simulation, which is why a wrong event is caught before it can be described
- arXiv. VideoPhy: Evaluating Physical Commonsense for Video Generation - one of the academic benchmarks for this problem. The search results for physics in generated video are almost entirely papers like this one, which is why we wrote the working version
- IEEE Spectrum. The Uncanny Valley (Masahiro Mori, translated by MacDorman and Kageki) - why a near-miss costs more than an obvious miss
Every figure above links to the place it was published. Numbers marked as ours are measured inside this studio and we say so where they appear. We do not print a statistic we cannot point at.
Send a link. Get one finished ad back.
One finished cut built from your own product, inside three days, free and yours to run whether or not we ever work together. Pause it on contact and run the eleven tests. If it fails one, you will know before we do.
Replies within a day. Ad within three.