Sutra

Duration is the action's pace, not a length

Asking a video model for seven seconds of a two-and-a-half second beat does not buy you spare footage. It directs slow motion. Here is what a duration field actually does, what to ask for instead, and why a tempo fix that saves one model ruins another.

What is in here
  1. How long should a shot be in an AI video?
  2. You are not buying spare footage
  3. How do you get more material inside a take?
  4. A tempo correction that saves one model ruins another
  5. When does a long take beat a cut?
  6. Two shot-length shapes that both beat an even rhythm
The short answer

As long as the action actually takes, and not one second longer. A generator's duration field is not a buffer, it is a pace instruction: you hand the model one action and a window, and it performs that action across the whole window. Ask for seven seconds of a two-and-a-half second beat and you have directed slow motion at roughly a third of natural speed. There is nothing else in the take to cut to.

What you get out of this
  1. Why a longer request returns slow motion instead of options, with the arithmetic that predicts it
  2. How to get real surplus inside a take: scene cuts in one call, not a longer single action
  3. Why a tempo correction is measured per model, and what happens when you apply one globally
  4. The shot-length cap we ran, and why we retired it
  5. Two shot-length shapes that both beat an even rhythm, with the frame counts

What a safety margin cost on a single film

One run, counted afterwards
2.5sthe length of the beat we actually neededSutra Haus production log
6-10swhat I requested per shot instead, as a margin against failureSutra Haus production log
90sgenerated footage bought to assemble a 25-second film on that runSutra Haus production log
2xthe ratio we hold to now: about fifty generated seconds for a twenty-five second filmSutra Haus production log
All four are ours, off one production run. The margin was insurance against having to come back with a failed shot. It protected the process and destroyed the footage.

01How long should a shot be in an AI video?

Time the beat first, then request exactly that. If the action is 2.5 seconds and the model's floor is 4, ask for 4 and accept a slightly slow take. Never round up to be safe. The duration you request is the pace you are directing, and the model has no concept of spare footage at either end.

I learned this the expensive way on a hotel shoot. I sized every shot with a per-shot safety margin: six seconds for a two-to-three second beat, ten seconds for a five-second hold, on the reasoning that a longer window gives you more to choose from. Ninety seconds of generated footage came back for a twenty-five second film, and nearly every take was in slow motion.

You're giving it the action once, and you're giving it seven seconds. So the model is making it really slow, which makes the videos unusable.

Badal Kariwal, reviewing that run

The same 2.5-second beat, requested at four window lengths

The arithmetic, not a measurement
Asked for 4 seconds63%Asked for 5 seconds50%Asked for 7 seconds36%Asked for 10 seconds25%
See the numbers as a table
What you asked the model forSpeed of the action, as a share of natural
Asked for 4 seconds63%
Asked for 5 seconds50%
Asked for 7 seconds36%
Asked for 10 seconds25%
This is division, not a study. The model performs one action across the window you give it, so the playback speed is the beat divided by the window. We publish it as arithmetic because that is what it is: the point is that the relationship is predictable before you spend anything.

Look at the bottom bar. A quarter speed is a person moving through treacle, and it fails worse than an obviously broken render, because the still frame looks fine. That is the class of error a frame check never catches.

02You are not buying spare footage

A longer take gives you options on a set: the actor does the thing, does it again, you pick. A generator does not work that way. It will not perform the action three times inside seven seconds so you can choose the good 2.3. It stretches the one action to fill.

What a longer window actually buys

What you expect, what you get
What a longer window actually buys
DimensionWhat a longer take gives you on a setWhat it gives you from a model
More attempts at the beatYes. Three or four takes to pick fromNo. One take, played slower
Handles either side of the actionYes. Room to breathe before and afterNo. No handles. The action starts at frame zero and ends at the last frame
Insurance against a failed shotYes. Real, and worth paying forNo. The insurance is what causes the failure
Something to trim toYes. Trim to the best 2.3 secondsNo. Trimming a stretched action gives you a shorter stretched action
What it costs when it goes wrongA slate and another takeNo. The window, again, at the correct length. A bad take is re-rolled, never repaired
The right-hand column is not a complaint about model quality. It is what the duration parameter means. Nothing in the tool is behaving badly.

That last row is a rule in this studio and it is worth stating on its own, because it changes how you budget: a bad generated take is re-rolled, never repaired. Speed-ramping a slow take back up strobes it. Cutting the middle out of it produces a jump in a shot that had no cut in it. The take is the unit.

03How do you get more material inside a take?

Ask for scene cuts inside the take rather than a longer single action. One call, three or four internal scenes, each with its own timeline slice written out in the prompt, each grouped by place so every segment stays anchored to something real. You get four usable beats at natural pace out of one window instead of one unusable beat stretched across it.

How we write a multi-scene call

The shape, not the words
ONE CALL, 10s, one location, three scenes

[0-3.5s]  hands enter frame, lift the jar off the shelf
[3.5-6.5s] quarter turn, label comes round to camera
[6.5-10s]  set down on the counter, hands leave frame

rules that make it work:
  every slice is a real action at its own natural speed
  every slice stays in ONE place, so nothing has to be invented
  no slice is longer than the action it names
  the word slow appears nowhere unless slowness is the point
The slices are the whole trick. Without them the model reads a ten-second window as one instruction and spreads whatever you asked for across all of it.

One window, spent two ways

The same ten seconds, twice
0s1s2s3s4s5s6s7s8s9s10sOne 2.5s action, stretched across the windowScene 1: hands enter, liftScene 2: the turnScene 3: set down, exitONE ACTIONTHREE SCENES
Read it as a list
AtChannelWhat happens
0.0sOne actionOne 2.5s action, stretched across the window
0.0sThree scenesScene 1: hands enter, lift
3.5sThree scenesScene 2: the turn
6.5sThree scenesScene 3: set down, exit
Top row: one beat, quarter speed, nothing to cut to. Bottom row: the same ten seconds and the same spend, returning three cuttable beats at natural pace. The difference is entirely in how the call was written.

Slices work because they take back a decision the model was making on your behalf. A ten-second window with one action in it is under-specified, and the model resolves it the only way it can, by slowing down. Three timed actions in the same window leave nothing for it to decide.

One call or three scenes?

Answer two questions
Every ending here is about the same thing: making the requested window match the action inside it. There is no branch that recommends a margin, because a margin is the failure.
Put this line at the top of the sheetBefore anything generates: target runtime, number of calls, total generated seconds, and the ratio between them, on one line. A twenty-five second film should read as roughly fifty generated seconds. When ours read ninety, the number was on no sheet anywhere, which is why nobody caught it until the footage came back.

04A tempo correction that saves one model ruins another

Generated motion often reads slow at native speed, so it is normal to nudge a clip faster before it enters a cut. We ran a standing 1.1 times correction on everything. It was right for the models it was written against, the ones that render people moving slightly underwater. On a model without that problem, the same nudge makes hands twitch.

Write the number down, with the model it was measured on

So the correction became a per-model number instead of a house constant. Measure the model's native pace on a clip whose real duration you know, apply a correction only where the footage genuinely reads slow, and write down what you applied in the build notes. That last part sounds like admin and is the part that matters: an undocumented multiplier surfaces the day somebody swaps models, when nobody remembers it exists.

Four pacing habits worth arguing with

Flip them
Every one of these is a position we held. The first two we retired. The last two we still hold, and the reasons are narrower than they look.

05When does a long take beat a cut?

When the shot has an arc of its own: something starts and finishes inside the frame, or the take carries its own scene cuts. That is the only test, and you apply it to the shot itself, watched cold at 360 pixels, which is the real width of a feed tile. Not to its description in a document.

A pour is the honest test for this. The liquid tells you the speed, so a pour at 40 percent of natural is not stylish, it is wrong, and every viewer knows it without being able to say why.

Liquid, hair, fabric and breath price a mistake instantly, because each carries its own physical clock. If one of them is in your beat, pace is not a taste question. They are also the events current models break most often, and that is not a coincidence: the clock that makes a pour readable is what makes it hard to fake.

What you should actually be requesting

Put your numbers in
Speed you just directed-
Request this instead-
Scenes to write into the call-
The speed figure is the beat divided by the window, which is what the model does with a single action. The scene count is the same arithmetic run backwards: if the floor is longer than your beat, that surplus has to become slices in the call rather than slowness in the take.

06Two shot-length shapes that both beat an even rhythm

Even is the enemy. Twenty rigorous films built to one cut rate read as a single film, and that sameness is what a viewer notices first and a maker notices last. Two shapes beat it, and they are opposites. One accelerates. One decelerates. Both are a direction, which an even rhythm never is.

A deceleration film, shot by shot

Frames per shot, in order
7f16f217f39f410f511f610f75f85f97f108f1112f1219f1332f1470f15
See the numbers as a table
Shot numberFrames
17f
26f
317f
49f
510f
611f
710f
85f
95f
107f
118f
1212f
1319f
1432f
1570f
Measured off one delivered film at 24fps. It opens in blocks of six to seventeen frames, tightens to a run of fives in the middle, then opens out to 19, 32 and 70 as the payoff lands. The last shot is nearly three seconds and it is earned by the twelve shots in front of it.

The acceleration shape runs the other way: three cuts across six seconds to set a pulse, then four cuts of 8, 14, 11 and 9 frames inside 1.4 seconds. And there is a third move worth stealing, for a sequence of similar shots that would otherwise read as an assortment. Halve the cut rate exactly halfway through: 45, 30, 30, 30, then 15, 15, 15, 15. The shots become a series because the rhythm is going somewhere.

Before you send the next batch of generation calls

Tick as you go - it remembers
0%
Six checks, about ten minutes, all of them before anything is spent. Every one of them exists because we skipped it once.

Questions people actually ask

Open what you need
How long should each clip be in a reel?

Long enough for one action and no longer. In practice that is 0.4 to 0.8 seconds for most beats, rising to 2.5 for a shot with an internal arc and 3.0 for a card carrying text. If you are generating rather than filming, the number matters twice, because it also sets the speed the action plays at.

Why does my AI video look like slow motion?

Almost always because the requested duration is longer than the action you described. The model has one action and a window to fill, so it fills the window.

Check the prompt for speed words too. Slow, gentle, unhurried, deliberate and graceful are all pace instructions, and stacking one of those on top of an over-long window compounds the problem rather than adding atmosphere.

Should I generate longer clips so I have more options?

No. A longer clip is one option played slowly, not several options. If you want real optionality, spend the same seconds on more calls with different actions. Motion's benchmark work across more than 550,000 ads found the top accounts and the average accounts share roughly the same 5 percent hit rate, with the top performers winning by shipping about 31 new creatives a week against 11. The lever is distinct swings, not spare seconds inside one.

Can I speed up a slow generated clip in the edit?

Sometimes, and it is worth trying once before you re-roll. Past about 1.3 times, frame blending starts to smear and motion strobes, and hair, liquid and fabric give it away first. A take that came back at a quarter speed cannot be rescued at all. The rule we run is that a bad take is re-rolled and never repaired, because repair time reliably costs more than the second call.

What is the minimum shot length that still reads?

About a fifth of a second, and only when the viewer already has context. A word per cut works at 0.2 seconds when the sentence stays on screen and the reader is picking up one new word each time. A word that replaces what was there needs roughly 0.75 seconds. The same distinction applies to pictures: a new subject needs longer than a variation on the one you just saw.

Does a fast cut rate hurt comprehension?

Not if the cuts carry grammar. A chain of unrelated clips at 0.2 seconds is noise; the same rate becomes readable the moment each cut is a word in one sentence. Cut rate is a symptom of structure. When fast cutting stops working, the fix is upstream in how the film is put together, not in the cut rate.

Every one of these mistakes is a margin somebody added to protect themselves. Longer windows, safer durations, a correction applied to everything so nothing gets missed. The margin always moves the cost from the person adding it to the film.

Where the numbers came from

  1. Motion. Creative Benchmarks, 550,000+ ads across 6,000+ advertisers and about $1.3bn of spend - used for the winner-rate and creative-volume figures in the FAQ

Every figure above links to the place it was published. Numbers marked as ours are measured inside this studio and we say so where they appear. We do not print a statistic we cannot point at.

Badal Kariwal

Runs Sutra Haus, a one-person ad studio that has shipped over a thousand finished creatives - film and stills - for DTC brands and hotels. Writes here about what the work actually taught him, including the parts that failed. The person who reads your brief is the person who builds the work. Send him something to make.

Bring us a beat you cannot get out of a model

Send a link. Get one finished cut back.

One 9:16 film built from your own product, inside three days, free and yours whether or not we ever work together. Watch the pours, the hands and the fabric. Those are where pace lies, and they are the first thing we would check in your current ads.

Replies within a day. Ad within three.
Read next