Sutra

Hands, products, and the contact problem

A hand got easy. A hand holding your product did not. Why a model improvises an object it has never seen, why the size is the first thing to go, and the five routes to a usable shot ranked by what each one costs you.

What is in here
  1. Why is contact harder than anatomy?
  2. Your product gets improvised, and the size goes first
  3. What if the product doesn't fit in the character's hand?
  4. The real-hand insert, and the shadow that gives it away
  5. How do you make AI UGC look real?
  6. When should you stop and shoot twenty seconds yourself?
  7. Hands are where our re-roll rate is worst
The short answer

A hand alone has millions of examples behind it. A hand on your product has none, because the model has never seen your product. So contact is a joint problem: two objects that have to agree about size, shape, pressure and shadow at the same time, and the model has only ever met one of them. It improvises a bottle-shaped thing, at the size bottles usually are, held the way bottles are usually held. Every fix that actually ships is a way of not asking.

What you get out of this
  1. Why grip and scale still break after anatomy stopped breaking
  2. Five routes to a product in a hand, ranked by what each one costs you
  3. How a real hand shot on a phone gets composited, and the shadow rule that stops it floating
  4. The point at which twenty seconds of your own footage beats any prompt
A woman holding a white facial device with a chrome ring and a black gloss panel against her cheekbone, fingers wrapped around the handle, with tick-marked benefit captions across the lower third of the frame
A still we made for a beauty device client. The wordmark printed on the device body is ours, not theirs. The ad rests on a few square centimeters: where her fingers close on the handle, and where the device stops being in front of her face and starts being against it.

01Why is contact harder than anatomy?

Anatomy is one object. Contact is two, and the second one is yours. Fingers stopped being a joke faster than anyone expected because the world has photographed hands endlessly. Nobody has photographed your pack, so the two have to agree about size, shape, pressure and shadow while the model is guessing at half of what it is holding.

The giveaway is that generated contact is one-way. Real contact deforms both parties: a fingertip flattens against a hard pack, skin goes pale at the pressure points, a soft tube dents. In generated footage the hand does all the work and the object does none. Look only at the pixels where the two meet and you will be right about the shot inside a second.

Four numbers from our own log

Ours, counted off delivered files
8ftthe height a stone entry crest was rendered at in one of our hotel films. The real one is about five, and the owner walks past it dailySutra Haus production log
20sof phone footage, which has rescued more failed shots for us than any prompt change ever hasSutra Haus
1.6pxthe entire edge treatment on a composited cutout: a Gaussian feather, no keyline, nothing elseSutra Haus
43%of the films that clear every automated gate we run are still killed by a human watching them afterwardsSutra Haus production log
All four are ours. The first is a mistake we made, published because the correction is the useful part and the excuse would not be.

Four things people say about AI hands

Flip them
None of these is a strawman. All four were said to us by somebody paying for creative, and the second one was said by us before we knew better.

02Your product gets improvised, and the size goes first

A model has never seen your product. What it has seen is the category, so it draws the average member of that category, at the average size, held the average way. The failure is rarely the object itself. It is the scale. Serums arrive the size of a wine bottle. A compact arrives the size of a hand mirror. Everybody on the review call spots the color first and the size never.

Four of our own frames, and what each one is asking a viewer to believe

Ours, made this way
None of these is a failure. They are the frames where a failure would have been obvious, which is why they were the ones watched hardest before they shipped.

We started writing sizes down after a hotel film rendered a stone entry crest at roughly eight feet when the real one is about five. Every shot naming a physical thing now carries its measured size beside the person's height, marked measured, owner-stated or assumed, and scale anchoring earns a piece of its own.

03What if the product doesn't fit in the character's hand?

Then stop trying to make it fit and change the shot. A product that reads wrong in a hand is a framing problem with five known answers, and four of them are cheaper than another generation. The order we try them in: cut away from the hand, frame the hand out, drop a real photograph of the pack in and move it as a card, composite a real hand, shoot the thing. What a still is allowed to do once it starts moving is set out in the physics test battery.

Five routes to a product in a hand

Ranked by what it costs
Five routes to a product in a hand
DimensionWhat it costs youWhen it is the right call
Generate it againA credit and twenty minutes, repeatedlyAlmost never. The same prompt has the same blind spot
Cut away from the handOne shot out of your editWhen the grip is not the point and the product is already established
Frame the hand outA crop, plus coverage you have to haveWhen you need the person and the mood, not the demonstration
Real product photo, moved as a cardOne photograph you already ownWhenever your own label has to be readable on screen
Real hand on a phone, compositedTwenty seconds of shooting, an hour of workWhen the product has to be picked up, opened, pressed or applied
Ours, in the order we actually try them. The first row is on the list because it is what everybody does first, not because it works.

04The real-hand insert, and the shadow that gives it away

This is the route people skip because it sounds like a visual effects job. It is twenty seconds of phone footage, a mask, and one shadow built correctly. The shadow is where almost every attempt falls down, for a single reason most people never notice they chose.

Compositing a real hand into a generated plate

Walk it
Step 1
Shoot the contact, not the product

Twenty seconds. Phone propped on a stack of books, one window as the only light. Pick the product up, hold it still for three seconds, put it down.

That is the entire shot list. You are capturing a grip and a release, not a performance, which is why it does not matter who is holding it or whether they can act.

Step 2
Match the light before you match the color

Find the direction, the height and the hardness of the key light in the generated plate. Then move your window, or turn your hand relative to it, until the shadows across your own knuckles fall the same way.

Color is a grade and can be fixed in a minute. A wrong light direction cannot be fixed at all, and it is the thing a viewer reads first without knowing they read it.

Step 3
Cut with one feather and nothing else

Our entire edge treatment on a composited element is a 1.6 pixel Gaussian feather. No keyline, no glow, no torn edge, no chromatic fringe.

Anything more decorative than that is a tell, because an edge in a real photograph is a property of the lens rather than a choice made by whoever cut it out.

Step 4
Build the shadow out of the element's own alpha

Not a drawn ellipse. Take the element's own silhouette, offset it, blur it, and multiply it down onto the plate. Where two shadows overlap, take the darker of the two rather than adding them, because light does not add darkness.

And express the offset as a fraction of the element's own width, never as a number of pixels. That one choice is the difference between a composite that sits and a composite that hovers.

Step 5
Watch it cold, at feed size

At 360 pixels wide most of what you fussed over vanishes and exactly one thing survives: whether the object is sitting on the surface or floating above it.

If it reads at that size, it reads. If it only reads at full resolution, you have made a print ad and put it in a video.

1 / 5

This is the method, not a parameter table. The numbers that matter here are the two we can publish: the feather, and the fact that the shadow offset is a fraction rather than a pixel count.
Why pasted-on composites read as pasted onThe offset between an object and its contact shadow scales with the object's own width. Copy a shadow recipe from a lipstick onto a suitcase without rescaling it and the suitcase floats, even though every number in the file is the one that worked yesterday. That is the whole thing: fractions of width, never a fixed number of pixels, and multiple shadows combined by taking the darker rather than by stacking them.

05How do you make AI UGC look real?

Make the two hardest things in the frame real: the contact, and the voice. Everything else can be generated without anybody noticing. In practice that means a real photograph wherever your label has to be read, a real hand wherever the product is handled, and an unscripted delivery with gaps in it, because the creator register is set by audio far more than by picture.

One of ours, in the creator register. Press play for the shot this page is about: the glass she picks up is the only object in the film that has to be gripped, and it is the frame we watched hardest before it went out.

A prompt built for contact rather than for looks

Build one

Assemble one and read the reasons. The pattern worth taking away is in the scale slot: a relationship generates better than a measurement. Ask for five feet and you get a five-foot-shaped guess.

One thing to steal from that scale slot, because it generalizes past hands. State the relationship, not the number. Reads about chest-to-eye height beside her produces a better result than five feet tall, because a model has no ruler and does have a very good idea of what a person is. The same instinct governs the rest of the sheet: describe what the frame should contain, not what the object measures. The failure modes behind all of this are cataloged in the frame-by-frame atlas of why AI ads look fake.

06When should you stop and shoot twenty seconds yourself?

The moment you have paid for the same shot twice. A second failed generation of a grip is not bad luck, it is the same blind spot returning, and a third will return it again. Twenty seconds on a phone, in one window's light, costs less than the credits already spent and produces something no prompt can: your actual product, at its actual size, in a real hand. If it turns out you need a face and a voice as well as a hand, that is a larger decision, and when you should still hire a human creator is the honest version of it.

Twenty seconds, and where to spend them

Move the sliders
The product entering frame
The grip, held still
The action: peel, press, pour, open
The release, and the hand leaving
Total

The default split is the one we ask brands for. It is weighted toward the action and the release because those are the two things a generated plate will not give you, and away from the entry, which almost any coverage can cover.

The twenty-second shot list, as we send it to clients

Take it
TWENTY SECONDS. ONE WINDOW. NO CREW.

Phone propped on a stack of books, lens level with the product.
One window on one side. Turn every other light off.
Wipe the product first. Fingerprints are the whole shot.

00-04   Hand enters. Pick the product up.
04-07   Hold still. Three full seconds. Do not adjust.
07-15   The action: peel, press, pour, twist, open.
15-20   Set it down. Hand leaves frame. Keep rolling.

Shoot it twice. Send both files.
No captions, no music, no filter, no crop.
Paste this to whoever is nearest the product. It has never once come back unusable, and it is the only brief in our whole process that does not need a checkpoint.

And the footage the brand already has

A client's own clips are the fastest way to solve contact and the fastest way to make an ad that looks like their last one. Borrowed material is never used raw here. It gets recropped, regraded and recomposed, and their captions and phone audio never survive. Borrowed footage is punctuation, never the spine, because the transformation is the only part anybody is paying for. If every hard shot is answered by a clip they already own, you have edited their camera roll.

07Hands are where our re-roll rate is worst

We have no clever answer for that. What we have is a rule that stops the bleeding: a bad take is re-rolled, never repaired, and every re-roll is somebody's decision before it is anybody's spend. Nursing a broken grip through four rounds of fixes is how a cheap shot becomes an expensive one. The economics are laid out in re-roll, never repair.

Four of our films, and the contact in each

Ours, made this way
Ephoria - campaign ad
Beauty device - campaign ad
Studio demo - a pour
Studio demo - skin
Every one of these mixes real product photography with generated coverage. Tap any of them to watch it full size with sound, and pause on the frame where a hand meets something.

The other honest limit: no film in our reference set carries a hook rate or a return on spend, so we can tell you which contact shots survive a human watching them and not that a better grip sold anything. Some categories are simply easier on this axis, and what kinds of product AI UGC actually works for is a better question than how to force the hard ones.

Questions people actually ask

Open what you need
Why do hands look weird in AI ad video?

Usually it is not the hand any more. It is the contact. Anatomy improved quickly because a hand is one object with an enormous amount of reference behind it, while a hand holding your specific product is two objects that have to agree about size, pressure and shadow at once. Watch the pixels where the fingers meet the object rather than the fingers themselves.

What if the product doesn't fit in the character's hand?

Change the shot rather than the prompt. Cut away from the hand, frame the hand out of the shot, or drop a real photograph of the pack in and move it as a card.

If the product genuinely has to be picked up or applied on camera, shoot twenty seconds of a real hand on a phone and composite it. That is cheaper than the third failed generation, and it is the only route that gets your actual label on screen.

How do I make AI UGC look real?

Make the contact and the voice real and let everything else be generated. A real photograph wherever the label has to be read, a real hand wherever the product is handled, an unscripted delivery with gaps in it, and a camera in a position a person could have held. Polish belongs to the object; rawness belongs to the performance. Mixing those up is what produces the look people call AI slop.

Can I just use the brand's own UGC clips?

As punctuation, yes. As the spine of the film, no. Borrowed footage gets recropped, regraded and recomposed, and the original captions and phone audio never survive. If a client's existing clips are answering every hard shot, the deliverable has quietly become an edit of their camera roll, and they can tell.

Is a still photograph of the product better than a generated one?

For anything where the label matters, always. A model has never read your packaging, so it writes something packaging-shaped, and your own customers are the people most likely to notice. Move the real photograph as a card - push, settle, drift, a light patch traveling across it - and never ask it to perform an event.

How long does the real-hand composite take?

Twenty seconds of shooting and about an hour of work, most of it spent on the shadow rather than the mask. The mask is the boring part. The shadow is where a composite either sits on the surface or hovers over it, and the deciding factor is whether the offset scales with the object's width.

Some products are held in ways nobody has ever filmed, and for those the honest answer is a camera and an afternoon, not a better prompt. Knowing which of your shots is that shot is most of the skill. Everything else on this page is refusing to ask a machine about an object it has never touched.

Where the numbers came from

  1. IEEE Spectrum. The Uncanny Valley (Masahiro Mori, translated by MacDorman and Kageki) - why an almost-right hand costs more than an obviously drawn one
  2. arXiv. VideoPhy: Evaluating Physical Commonsense for Video Generation - the academic framing of the same problem: models are scored on appearance far more reliably than on interaction

Every figure above links to the place it was published. Numbers marked as ours are measured inside this studio and we say so where they appear. We do not print a statistic we cannot point at.

Badal Kariwal

Runs Sutra Haus, a one-person ad studio that has shipped over a thousand finished creatives - film and stills - for DTC brands and hotels. Writes here about what the work actually taught him, including the parts that failed. The person who reads your brief is the person who builds the work. Send him something to make.

The contact shot, made for your product

Send a link. Get one finished ad back.

One finished cut built from your own product, inside three days, free and yours to run whether or not we ever work together. Pause it on the frame where a hand touches the thing. That is the frame we would look at first.

Replies within a day. Ad within three.
Read next