What these models can and cannot do
Two camera tests, run against the same video models
StoryBarn generates with. Not a demo reel — the point is where they fail, because
that is what decides whether a series holds together, and it is what nobody shows you.
The claim: the same person twice
Every tool can make one good-looking clip. The question for a series is
whether the second shot has the same person in it. These are two separately generated
shots, both animated from one approved reference, cut together — no retouching.
Cap, beard, jacket, tool belt and build all carry
across the cut. That is what a locked reference image buys, and it is the whole basis of
everything else here.
Play it with the sound on. Nothing asked for audio and it came back
anyway — wind across an open yard, boots on frozen ground, a jacket creaking. Both
halves were generated separately, so the room tone had no reason to match across the
cut, and it does.
What money buys, and what it does not
Three people outside a barn, one speaking and gesturing, the others
listening — an ordinary shot, and what a drama is mostly made of. Same prompt,
same reference, two price tiers.
The reference is itself generated, from three
locked cast references — the tool's own pipeline tested on its own output, which
is what making a show actually looks like.
The cheaper tier.
Four times the price.
Both hold. The hands separate them.
Every face survives the full eight seconds on both, in the right clothes
and the right places, and neither cut to a shot nobody asked for. The difference is
finger-level: the gesturing hand at three seconds, cheaper tier above.
One renders a soft mitten — the shape of a
hand without fingers. The other separates them. At four times the price, that is the
whole of what the money buys on a shot like this.
The sound is real, and mostly unusable
These clips generate their own audio, synced to their own picture. On the
barn shot that is a gift: wind and footsteps land on the frame that made them, which is
the one thing a library of sound files can never do.
On people talking it is a trap, and the reason is worth knowing before you
pay anybody for AI video. The mouths are moving to that audio — and a series
has to use its own cast, chosen once and held for thirty episodes. StoryBarn replaces the
generated voices with yours, which leaves lips synced to a track nobody hears. Keeping
the generated voice instead is no escape: it cannot be named or kept, so every clip is a
different person.
This is why the tool checks the work instead of trusting it. A twenty-minute
episode is around 165 of these clips. Nobody watches 165 of them hunting a second and
a half of something wrong — so StoryBarn looks for you, says which shot and at
what second, and will not lock a reference that failed.
Every generator has these failure modes. What you should ask
of any of them is what happens next.
Where each tier belongs
| When you are | What it costs | What you get |
| Blocking a scene — timing, staging, where the cut falls |
cheapest |
Real motion, and a different person every clip. Fine, because identity is not
the question yet |
| One or two people, talking | cheap |
Indistinguishable from the top tier in our tests |
| Two or three people, talking | cheap |
Both hold every face; the gap is finger detail nobody is watching for |
| Hands or an instrument as the subject | full |
Worth the four times when that is what the shot is of |
Tested 2026-09-10 on eight-second clips at 720p,
one generation at a time. Videos here are re-encoded to load quickly; they carry the
audio they came back with.