42 min readMarkdown
Grok Imagine Video 1.5: the complete prompting guide
One picture in, up to fifteen seconds out, with sound you cannot turn off. Nine rules, three pipelines and the failures worth knowing about first.
Grok Imagine Video 1.5 animates a picture you give it. One to fifteen seconds, sound made in the same pass, and no way to run it from words alone: every clip starts from an image, and that image comes back as frame one almost exactly.
That is the whole shape of the model. There is no reference slot, no last-frame slot and no audio switch, so most of what would be a setting on another model is either already in your picture or not available at all.
The picture is the brief. The prompt is only what changes.
It also builds the clip front to back: frame one first, then each moment from the one before it, rather than planning the whole thing and filling it in. Half the rules below come from that — why your opening composition survives, why the order of your sentences is the order of the film, and why the model would rather turn a head than cut.
Overview
The short version
Four decisions do most of the work. If you read nothing else, read these. Each one is written out in full further down.
-
Spend your effort on the picture, then describe only what moves
Your image comes back as frame one almost untouched: 35.5 dB against the original, and three takes on three unrelated pictures agreed to within 0.15 dB. Framing, wardrobe, props and light all survive. Describing them again wins nothing, and invites the model to redraw something that was already right.
Instead of
A raccoon on a doormat at night, infrared, a white door behind it…
Write
The frame is already correct. Animate it and change nothing else.
-
Telling it what not to do does nothing
xAI say plainly that negative phrasing is ignored, and our runs agree. One fifteen-second brief demanded hard cuts and banned dissolves eight times over, and came back with a cross-fade at all nine of its shot changes. Ask for what you want as something that happens, never as something that must not.
Instead of
No camera movement, no zoom, no push-in, no drift
Write
The framing at ten seconds is identical to the framing at zero seconds — same crop, same angle, same lens
-
You cannot ask for a cut. You can make one necessary
Grok changes shot the way a body would. Ask for a new angle at the same distance and it turns a head. Ask for a new distance and no movement of the body can do that, so it cuts. One take with no cut instruction anywhere came back with two clean hard cuts, both of them size jumps.
Want an unbroken take? Pin the shot size — she is the same size in the frame in the last second as in the first. Want cuts? Write shots at different distances, or render them separately and join them afterwards.
-
The sound answers events, and there is no off switch
Every take makes audio whether you write for it or not. What sets the level is not your wording but how many things happen. Three takes asking only for room tone came back between A loudness measurement. A finished social video sits around −16 LUFS; −70 LUFS is silence with a file size. — inaudible — while a take naming a klaxon, footfalls, rain and a shutter came back at −15.8.
Write the sound as a list of events, each with a material and a surface, not as a mood. If nothing in the shot makes a noise, expect a silent file and plan to lay your own track underneath.
What it accepts
| Input | Limit | Notes |
|---|---|---|
| Prompt | 4,000 chars | Roomy. Your picture is the limit here, not the character count. |
| Starting image | 1, required | It becomes frame one. There is no text-only door. |
| Duration | 1–15 seconds | Whole seconds, anywhere in the range. No automatic option. |
| Resolution | 480p · 720p | Same price, so there is no reason to shoot the smaller one. |
| Aspect ratio | Auto + seven | 16:9, 9:16, 1:1, 4:3, 3:4, 3:2, 2:3. Auto means the picture’s own shape. |
| Reference images | none | No reference channel exists. Identity comes from frame one or from nowhere. |
| Last frame | none | You cannot pin the ending to a picture. |
| Audio | always on | No switch, no silent mode, no separate charge. |
Seven named ratios is the widest set here after Seedance, and one second is a real length rather than a rounding error — useful as a punctuating beat in an edit. Against Veo 3.1 and Seedance 2.5 it gives up every other input: no references, no last frame, no silence.
What it costs
One flat rate, and the simplest price in the catalog: about 6.7 credits a second, whatever the resolution and whatever the shape.
| Length | Cost | What it’s for |
|---|---|---|
| 1s | 7 | A punctuating beat, or the cheapest look at how a frame moves |
| 5s | 33.5 | The steady length: one subject, one action, one camera idea |
| 10s | 67 | A complete short piece — a vlog, a found-footage gag, a product spot |
| 15s | 100 | The ceiling, and the length most likely to drift |
480p and 720p cost the same, so a cheap test is a shorter clip, not a smaller one. On Seedance you save by dropping the resolution. Here you save by halving the length.
A one-second take at 7 answers a lot of questions: whether the picture is accepted, whether the ratio came out the way you set it, whether the model reads the shot the way you do.
The rules
Nine things that are true whatever you are making. Each one came out of a take that cost money, and they are ordered by how much they will save you.
Describe what changes, not what the picture already shows
The image is not a hint. It is frame one.
Measured against the still we handed in, the clip’s opening frame comes back at 35.5, 35.6 and 35.65 How close two pictures are, measured as peak signal-to-noise ratio. Above about 35 dB the difference is hard to see at all. Around 20 dB you are looking at a redrawing rather than a copy. across three completely different pictures — a character sprinting down a server aisle, a bedroom selfie, a mirror shot. A model that treats a first frame as an anchor rather than a lock scores 21.5 dB on the same test. Fifteen decibels apart is a different tool.
So everything the picture settles is settled: framing, wardrobe, props, the state of the room, the light. The prompt’s only job is motion.
The lock is weaker in one place, and it is worth knowing before you build a run on it. A take whose subject wore an embroidered Mongolian deel came back at 32.46 dB overall, and the loss was not where you would guess:
| region | frame-one fidelity |
|---|---|
| her face | 34.04 dB |
| the grassland behind her | 33.35 dB |
| the deel, belt and embroidery | 27.71 dB |
The face is the best-preserved thing in the picture. What gets repainted is fabric. The belt’s medallion, its amber stone and both tassel pendants are the same objects with the same design ten seconds later — the threads are what get redrawn. If your subject is a garment, keep your detail at the level of the objects on it, and do not expect 2048-pixel embroidery to arrive intact at 720p.
Telling it what not to do does nothing
The model ignores negative phrasing. xAI documents this, and we paid to confirm it.
A fifteen-second action brief wrote HARD CUT. three times inside its beats,
banned dissolves and fades in its style line, then banned dissolve, fade and
wipe again in its Avoid: line. Eight instructions, zero hard cuts — and a
cross-fade at all nine shot changes. The same brief on Seedance cut eight times.
The fix is never a ninth ban. It is to find the positive sentence that makes the thing you do not want impossible:
- “no camera movement” → “the framing at ten seconds is identical to the framing at zero seconds — same crop, same angle, same lens, same distance from the mat”;
- “no change of shot size” → “her arm’s length never changes, so she is the same size in the frame in the last second as in the first”;
- “no readable text” → make the thing too small and too far out of focus to read, so the composition never reaches for a string in the first place.
Keep an Avoid: line at the foot of a brief for one reason only: a bad result
is then a result that failed against instruction, which is more useful to
know than one that failed in silence. Do not spend characters on it expecting
work in return.
You cannot ask for a cut. You can make one necessary
How a shot change is handled is decided by the content, not by the prompt.
This is the most useful thing to know about the model, and it took four takes to pin down.
- A ten-second bedroom vlog with no cut instruction anywhere came back with two clean hard cuts, at 1.83s and 3.62s.
- A fifteen-second chase demanding nine hard cuts came back with none.
- Two more ten-second takes — a mirror selfie, a ride on horseback — came back with zero cuts at every detection threshold from 0.4 down to 0.1.
The pattern has nothing to do with the words. It is about whether a body could have done it:
| what the next shot changes | what the model does |
|---|---|
| the angle, at the same distance | turns the neck, the hips, the whole body — no cut |
| the distance (medium selfie → extreme close-up) | cuts, because no movement of a body moves the lens |
Both cuts in that vlog were size jumps. Both of its non-cuts were angle changes, done as one continuous turn in front of a phone that never moved. The chase asked for nine shot changes that were all, as the model read them, somewhere a hand-held camera could have walked to — so it walked.
Two consequences. For an unbroken take, pin the shot size, and say it as
something that is true rather than something to avoid. And a piece that really
needs eight hard cuts is eight generations, rendered separately and
joined afterwards, not one brief
shouting HARD CUT.
The sound answers events, and there is no off switch
audio_mode: required. There will be a track. The only question is whether anything is on it.
Seven takes, measured the same way. Sort them by what their sound block actually named and the pattern is plain:
| what the brief’s sound block asked for | integrated loudness |
|---|---|
| room tone under a silent performer | −70.0 LUFS |
| room tone under a silent performer | −69.1 LUFS |
| room tone under a silent performer | −64.7 LUFS |
| night ambience plus a few small events | −48.7 LUFS |
| one sustained crowd bed | −38.1 LUFS |
| one vocal event — a laugh | −26.9 LUFS |
| a continuous stack: klaxon, rain, footfalls, a shutter | −15.8 LUFS |
The three at the top are silence with a file size. Asked for a hum with nothing to attach it to, the model gives you nothing at all rather than something quiet. Asked for a laugh, it gives a laugh at −26.9. Asked for a chase, it gives a full mix.
So write sound the way a sound designer would: name the event, the material and the space. “City sounds, traffic” is a mood. “Cars passing, skateboard wheels on pavement, a distant rumble down the street” is a list of things that happen.
Two smaller findings worth carrying:
A face turned away is a face with no track. In the ride, the laugh starts at 2.86s and drops twenty decibels back to bed level at 6.40s — the exact frame her head turns away from the lens. She is still visibly laughing until 8.5s, so two full seconds of it are silent. The simplest reading is that the model mixes a voice to the camera rather than to the character. If a sound has to be heard, play it to the lens.
Ban speech twice if nobody speaks. A smiling performer with no stated sound is an invitation to write her a line and lip-sync it in whatever language the model fancies. Say it in the sound block, then again as an action — “she never opens her mouth to speak”. That ban held in every run here, which makes it one of the two negatives worth writing, alongside “no music”.
Ask for the end state, never an amount
An amount reads as a gesture. A destination reads as a rule.
A broadcast recreation asked for a push-in of “no more than three percent across the whole clip, so slight it is nearly invisible” and got roughly twenty percent: by the fifth second the foreground extra had left the frame.
Written as a claim about the last frame, the same instruction is exact. A doorbell piece said: “the framing at ten seconds is identical to the framing at zero seconds — same crop, same angle, same lens, same distance from the mat.” Frame 0 and frame 239 put the door edge, both hinges, the panel lines, the mat edge, the step curve and the vignette on identical pixels.
This works beyond camera moves. Anything you can phrase as “X is the same at the end as at the beginning” is something the model can hold on to. Anything phrased as an amount is a suggestion.
A cut can land early; a body cannot
Beats keep their order. What moves them is how long a performance takes.
Timestamps are read, and the order of your beats is the order of the film. What varies is how far a mark slips, and that depends on what the mark is:
- Two shot changes that came out as cuts landed early: asked 2.0s → 1.83s, asked 4.0s → 3.62s.
- Four that came out as movements the subject performs all landed late: +0.4s, +1.0s, +2.5s, +1.0s. The +2.5s pushed a whole window along with it.
A cut takes no time and can go anywhere, including ahead of its mark. A 180° turn takes as long as a body takes, and everything queued behind it slides.
The confirmation is a take that scheduled one turn with nothing after it. Its two marks landed at −0.14s and −0.10s, both early, and the turn then ran two seconds long at no cost to the schedule, because nothing was waiting on it.
Budget a performed beat by how long the movement takes, not by how many beats fit in the runtime. Never put a turn in the middle of a full schedule.
One more thing to expect from a beat sheet. On a format the model knows well, a big familiar gesture arrives whether you asked for it or not, and a small specific one loses to it. A brief asking for “a quarter turn of the hips” got a full turn away from the mirror, held for three seconds — the standard move of that genre — while “glances down at the phone for a moment”, which is not part of that vocabulary, never happened at all. Treat the beat sheet as a schedule the model fills, not a script it follows.
Spend one action on the ending
It slows into the last second. Seedance speeds up out of it.
Over the final 1.4 seconds of the same closing beat, Grok ran 14.5 → 4.4 on frame-to-frame difference, slowing to a stop, while Seedance ran 7.0 → 17.6 and was at its fastest when it ended.
Neither ends on a frozen frame, but the consequence is real. That closing beat asked for two actions: “shoves the deck into her hoodie pocket and pushes up out of the crouch to run again.” Grok did the first and settled.
So the rule that a final block must name an event — a Veo and Seedance finding — does not carry over as written. Here the final block should name exactly one event, and it should be the one you want on screen at the buzzer.
It holds the picture, not the objects in it
How faithful the frame is and how stable an object is are two different things, and only one of them is 35 dB.
The chase’s first frame has the right prop: a flat matte-black handheld computer
with a hot-pink X on the lid. By the second second she is carrying a red-cased
smartphone with a pink X on its wallpaper, and it stays a phone through most of
the middle of the clip. A whole Prop: paragraph naming its shape, its colour
and its glowing edge strip did not stop that.
Text splits the same way. A broadcast frame’s graphics came back with every
string correct — the tournament name, the score, the clock, the network
wordmark, the LIVE flag — and the layout redrawn: smaller, cell outlines
gone, a ▶ grown beside LIVE. That new design then held steady for the rest
of the clip. A thinner overlay in another take, a corner timestamp and a REC
dot, survived untouched.
Read it this way: the model copies the picture you gave it, then rebuilds every object in it from whatever the moment before says. Anything that has to be exact from a new angle is not safe in the prompt. If a graphic has to be right, composite it afterwards, where the model never sees it.
It will animate a photographed face
The refusal you hit on Seedance is not here, and it is often the whole reason to be here.
Hand Seedance a photorealistic person and it declines: E005 — the input or output was flagged as sensitive, through the reference door and the first-frame
door alike, six of six on one run and seven of seven on another. Grok takes the
same kind of picture without complaint — and since it needs an image, there is
no other way to use it.
That makes it the second way around that limit, after Veo 3.1, and a wider one in one respect: Veo’s reference route forces 16:9, while Grok took a 4:3 frame and charged nothing extra for the shape.
It is not a licence. A face you have no right to use is still a face you have no right to use, and every person in the examples here was generated. The model’s own filter is also strict about named people, brands and logos, and it misfires on harmless prompts. Describe a fictional person and keep brands generic rather than arguing with it.
The pipelines
Grok has one door: an image and a prompt. What changes between jobs is where frame one comes from — and since it is your entire identity budget, that choice matters more than the prompt does.
The blocks every brief uses
Whichever way you got your picture, the prompt is built from the same parts, in the order the model reads most reliably. It renders front to back, so prompt order is timeline order: an action written at the top of the brief happens at the top of the clip, and one buried at the bottom may arrive late or never.
[Continuity] who and what must not drift, named once
[Camera] where the lens is, and the end state that proves it stayed
[Motion] what moves, and how hard
[Beats] the schedule, stamped, in the order it happens
[Sound] events, with their materials and their spaces
[Style] the stock, the grain, the frame rate
[Avoid] kept for the record, not for the result
There is no [Look] block, and that absence is the point: the light, the
colour, the wardrobe and the room all arrived with the picture. Writing them
again spends characters on a decision already made, and invites the model to
redraw it.
Cut anything the picture already shows. If a sentence would be true of the still, it does not belong in the clip’s brief.
Two more things about the beats. Keep them coarse — one subject, one action, one camera idea each, which is xAI’s own advice and matches what we measured. And use strong words for how hard things move, because without them the model picks its own reading and it is usually milder than you wanted: “car passing” becomes “car racing past at high speed”, “the wave crests” becomes “crests fully and pitches forward, crashing down”.
How to write the sound block
Sound is not optional, so this is never a block you get to skip. Two notations work, neither better than the other, and both beat burying a sound inside a sentence about the picture.
Sound: … a paragraph among the others
AUDIO: … a section at the very end of the prompt
What goes in it is events, each named with the material it happens on and the space it happens in.
Sound: night ambience, faint crickets and a distant road. Claws on concrete
and the scuff of cardboard as the first parcel is set down. A beat of near
silence as the eyes open. Then many sets of claws on concrete at once, and
cardboard settling on cardboard as the stack builds. No music, no voices.
Three things worth knowing before you write one.
Silence is not on the menu, but it is often what you get. If nothing in the shot makes a noise, expect a track at −65 LUFS or below and plan to score it yourself. It also means a clip you intend to publish silent costs nothing extra — the audio comes out of the same per-second rate either way.
Quote a line if you want a line. Short dialogue works, and lip-sync is decent from a front-facing portrait with the mouth visible. Keep the line short, keep the camera locked, and put the tone in the sound block (“hesitant, then determined”).
“No music” is one of the two negatives that seem to hold, the other being a ban on speech — but read that honestly. None of our takes came back scored, and none of them was a scene with music in it. On a piece that invites a soundtrack, expect an argument rather than an instruction.
Pipeline A: draw frame one
The default, and what most of the examples here do. Make the still on an image model, look at it, fix it, then hand the finished picture over.
What it is for
- Best for: anything invented — a set, a creature, a found-footage conceit, a piece of branded chrome, a character who exists in one file.
- Costs you: a second generation and a second look, usually between 4 and 11.
- Watch for: the still is your whole identity budget, so it has to contain a face if a person is in the piece. A detail shot hands the model a pair of hands and asks it to invent a woman for the next ten seconds.
How to write the still
Three things belong in frame one that people habitually leave for the clip:
- Staging. A messy bed, kicked-off shoes, a stack of clutter. A model asked mid-clip to mess up a tidy room will mess it up while the camera is running.
- Text. A video model asked to invent a headline gets close and no closer. An image model spells it. Put the wordmark, the timestamp and the scorebug in the picture and let the clip preserve them.
- Shot size. Frame one is the opening shot, at 35 dB. Compose it as the first shot of the film, not as a portrait of the subject.
And one thing that belongs nowhere: make unwanted text unreadable rather than banned. A wall of photo prints described as “too small and too far out of focus to read” leaked no text at all. A banned string the composition still wants is a string the model draws badly.
Worked examples: Nobody filmed this and Nine shots and not one cut.
Pipeline B: animate a picture you already have
A generation already in your library, or a photograph. No second render, no second charge — and frame one is fixed, rather than something you can buy your way out of.
What it is for
- Best for: the cheapest clip there is, since the picture is already paid for, and for bringing back something you made weeks ago.
- Costs you: the lever. On a still you draw, anything the clip needs can be added in step one. Here the ten seconds have to be written to the picture you already have.
- Watch for: the aspect ratio. Left on Auto, the clip takes the image’s own shape, which is how a 4:3 clip ends up in a 9:16 feed.
How to write it
Read the picture first, and write only what it can support. The ride below is the case: the frame holds a mane, a saddle, a rein and a slice of shoulder — no legs and no head — so the brief never describes a gait. It describes the evidence of one, and puts the gait itself in the camera: “a light regular bounce in time with the horse’s stride.”
Measuring the result says it worked. The frame moves up and down far more than side to side (2.3 px against 0.9 px per frame on average), the bounce holds a steady 1.6–1.8 Hz for the full ten seconds, and over 241 frames it drifts 28 px vertically on a 1280-pixel picture. A pan piles up; this does not.
A frame that does not contain the hard thing does not have to solve it.
Some pictures hand you a rule for free. A selfie is one: an arm cannot change length, so the shot size is pinned by physics rather than by a promise — which, by the cut rule above, is an unbroken take you did not have to argue for.
Worked example: A ride nobody filmed.
Pipeline C: several takes, cut in post
Fifteen seconds is the ceiling and hard cuts cannot be prompted, so anything longer, or anything genuinely edited, is more than one generation.
What it is for
- Best for: a montage, a sequence with real cuts, anything past fifteen seconds.
- Costs you: continuity, which no input here can carry. There is no reference channel and no last-frame slot.
- Watch for: drift between takes. This is where Grok is weakest against the alternatives, and it is worth saying so plainly.
How to chain
Extract the final frame of one take, hand it in as the next take’s starting image, and join the run at the end. Frame one is a 35 dB lock, so the seam is genuinely tight — tighter than the same trick on a model that treats a first frame as an anchor.
What the chain cannot do is come back to a subject after a shot they were not in. If your piece needs one character across eight shots from eight angles, the honest answers are a model with a reference slot — Seedance 2.5 takes thirty — or an image model holding a character sheet, with Grok animating each frame it produces.
We have no finished chain to show yet, and inventing one would be worth less than saying so.
When it breaks
“It will not let me generate — it wants an image”
There is no text-to-video door on this model. The composer blocks the request before it costs anything, with “Grok Imagine Video 1.5 needs an image of the subject — attach one before generating.”
Fix: make a still first, or pick one out of your library. If you wanted a clip from words alone, that is Seedance or Veo rather than this.
“My cuts did not happen”
You cannot set the transition style from the prompt. Every ban, every
HARD CUT., every “nine shots” in a style line — none of it changes how a
shot change is handled.
Fix: if the next shot is at a different distance from the one before it, the model cuts on its own. If it is not, no wording will buy you a cut, and the piece has to be rendered as separate takes and joined.
“The aspect ratio is not the one I asked for”
Left on Auto, the clip takes the starting image’s own shape.
Fix: set the ratio explicitly, and make the still in the shape you need so the two agree. A 4:3 picture handed to a 9:16 setting is a crop somebody has to decide about, and it is cheaper to decide it in step one.
“The audio track is silent”
The brief asked for atmosphere and gave the model nothing to attach it to. Ambience on its own comes back at −65 LUFS or below, which nobody can hear.
Fix: name events, with the material they happen on — footsteps on gravel, a mug set down on wood, a shutter crashing, a laugh. If the shot really has nothing in it that makes a noise, accept the silent file and lay your own track under it. The audio was not charged for separately anyway.
“The camera moved when I told it not to”
You gave it an amount. A percentage, a “slight”, a “nearly invisible” — each of them reads as a gesture toward movement, which is movement.
Fix: state the end state instead. “The framing at ten seconds is identical to the framing at zero seconds” produced a pixel-identical pair of frames over 240 frames.
“Everything after the first beat arrived late”
There is a performance in the way. A turn, a walk, any big physical move takes as long as a body takes, and every beat behind it slides.
Fix: put the turn last, or give it a window of its own with nothing queued behind it. Cuts can land early; bodies cannot.
“The last second goes slack”
Two actions in the closing beat. The model performs the first and settles into the end.
Fix: one action in the final block, and make it the one you want on screen at the buzzer.
“The object in my picture turned into a different object”
The model copies the frame you gave it, then rebuilds each object from scratch as the angle changes. A prop described in the prompt does not survive that on its own.
Fix: keep the object at a consistent angle, keep the take short, or render the shots that show it from a new angle as separate generations. If it is a graphic rather than a prop, composite it afterwards.
“The text in my frame came back redrawn”
The words survive; the design does not. Expect the spelling to be right and the layout to be reinterpreted, then held steady for the rest of the clip.
Fix: anything whose design has to be exact is a compositing job. Anything whose wording has to be exact belongs in the still, large in frame.
Worked examples
Three finished pieces, each with the brief that made it. Every one is a single generation — nothing here is several takes cut together, which on this model means nothing here has a cut in it unless the model chose to put one there.
Nobody filmed this
Pipeline A: a drawn plate, then ten seconds of a camera that cannot move.
- Length
- 10s, one take
- Frame
- 9:16, 720p
- Supplied
- One drawn plate
- Cost
- 11 + 67
The plate, made first — including the burned-in camera chrome.
A doorbell camera chooses nothing. It is bolted to a wall, it cannot pan, it cannot cut, and it does not know it is filming. That makes the fixed mount the whole genre, and one instruction matters more than all the others.
The camera did not move once. Frame 0 and frame 239 put the door edge, both hinges, the panel lines, the mat edge, the step curve and the vignette on identical pixels. That is the end-state rule at work: an earlier broadcast take asked for a push-in of “no more than three percent” and got twenty, while this one asked for the framing at ten seconds to equal the framing at zero.
The camera chrome is in the plate rather than the prompt, and it survived
untouched — FRONT DOOR, the timestamp, the battery and the REC dot, same
style, same positions, frozen. A thin corner burn is closer to what the model
expects than a broadcast graphic, and the broadcast graphic is the one that came
back redrawn.
Two things it teaches about sound. The eye-shine gag needed no sound at all, and
the model made its own atmosphere for it — but the brief’s Audio: line is
mostly ambience, and the track measures −48.7 LUFS. Audible if you go
looking, invisible in a feed. The chase below asked for a klaxon, rain,
footfalls, a shutter and a body hitting brick, and came back thirty decibels
louder.
The clip's brief, in full
The frame is already correct. Animate it and change nothing else.
Camera: bolted to the wall beside the door. It cannot move. The framing at
ten seconds is identical to the framing at zero seconds — same crop, same
angle, same lens, same distance from the mat. No pan, no tilt, no zoom, no
push-in, no handheld drift, no rack focus, no cut.
0-3s: the raccoon crouches and sets the white parcel down on the doormat,
then nudges it with one paw until it sits square to the edge of the mat.
3-4s: it rises back onto its hind legs and looks straight into the camera
lens, and holds still.
4-5s: behind it, in the solid black beyond the mat, pairs of small bright
eyes open and catch the infrared. A dozen of them, spread across the
darkness at the same low height. Nothing else of them is visible yet. The
raccoon in front does not turn around.
5-10s: they walk forward out of the black into the infrared light, one
raccoon behind each pair of eyes, every one of them carrying a plain white
parcel of its own. They reach the mat unhurried and set their parcels down,
building a neat stack. The first raccoon steps aside and lets them work. At
ten seconds the stack stands on the mat and several of them are looking
directly into the lens.
Light: constant infrared, blown out near the door, solid black past the mat.
No exposure shift, no flicker, no colour anywhere, no porch light switching
on.
Overlay: FRONT DOOR, the timestamp and the REC dot are burned into the
frame. They do not move, drift, warp or re-spell.
Audio: night ambience, faint crickets and a distant road. Claws on concrete
and the scuff of cardboard as the first parcel is set down. A beat of near
silence as the eyes open. Then many sets of claws on concrete at once, and
cardboard settling on cardboard as the stack builds. No music, no voices, no
dialogue.
Avoid: camera movement of any kind, cuts, dissolves, speed ramps, a raccoon
that morphs or changes size, a human appearing, the door opening or closing
further, colour appearing in the frame, on-screen text other than the
overlay already in the frame, captions, watermarks, a doorbell chime.And the plate it animates, made first on GPT Image 2. Every word the clip has to keep is written here, not in the clip's brief.
A frozen frame from a home doorbell camera at night, infrared night vision.
Look: monochrome with a cold blue-grey cast and no colour anywhere. A harsh
infrared floodlight from the camera itself blows out the doorframe and the
near concrete to almost white, and falls off to solid black about two feet
past the doormat. Wide fisheye lens, strong barrel distortion bending the
edge of the step into a curve, dark vignetted corners. Soft focus, visible
sensor noise, heavy compression, the flat cheap look of real security
footage.
Framing: vertical, the camera fixed to the wall beside the door at chest
height and angled slightly down onto a concrete stoop. On the left a white
panelled front door stands half open, its small window catching the
infrared. Centre, a dark rectangular doormat on the concrete. Beyond the
mat, total blackness.
Subject: a raccoon standing upright on its hind legs on the doormat, facing
the camera dead-on, holding a small plain white cardboard parcel against its
chest with both front paws. Thick fur, ringed tail, black mask across the
eyes, eyes bright and reflective in the infrared. Lightly stylised, a
character from a 3D animated film rather than a photograph of a wild animal.
Calm and deliberate, as if making a delivery.
Overlay burned into the frame, thin white sans-serif, crisp and correctly
spelled. Top left: FRONT DOOR. Top right: 08/05/2026 02:47:13. Bottom left:
a small battery icon, then a small filled dot, then REC.
No other text, no people, no other animals, nothing else on the porch.A ride nobody filmed
Pipeline B: a picture that already existed.
- Length
- 10s, one take
- Frame
- 9:16, 720p
- Supplied
- One existing picture
- Cost
- 0 + 67
Already in the library, already paid for. Nothing was uploaded.
The cheapest clip on this page, because frame one was made weeks earlier for something else. Sixty-seven credits, first submission, no re-rolls.
The picture decides almost everything. Her arm leaves the frame at the bottom left, so the camera is a hand on a moving horse — and of the horse, the frame holds a mane, a saddle, a rein and a slice of shoulder, no legs and no head. So the brief never asks for a gait. It asks for the mane to lift, the hair to stream, the tassels to swing and the ridge to slide past, and puts the trot itself in the camera as “a light regular bounce in time with the horse’s stride.” The bounce came back steady at 1.6–1.8 Hz for the full ten seconds.
The audio is why this one is worth watching with the sound on. Three earlier takes asked for room tone under a silent performer and came back silent. This one asked for a laugh, and the laugh arrives at −26.9 LUFS, forty decibels above them. It also stops dead at 6.40s, the frame her head turns away from the lens, while she is still visibly laughing for another two seconds.
And the thing nobody wrote: “looks off past the camera toward the hills for a moment” came back as a head thrown fully back into the wind, hair streaming, held for two seconds. That is the standard image of this format, and both halves of the instruction — toward the hills, for a moment — lost to it.
The brief, in full
Three beats in ten seconds, deliberately wide windows, and not one sentence describing the light, the clothes or the land she is already standing on.
The woman in the first frame is riding a horse at a steady trot across open
Mongolian grassland, filming herself on her phone held out at arm's length.
One continuous unbroken take, ten seconds, no cuts and no transitions of any
kind.
She is the same woman the whole way through — same face, same long dark
wind-blown hair, same dark brown deel with gold-brown embroidered trim, same
wide silver belt with the round silver medallion and its amber stone, same
pair of silver tassel pendants hanging at her hip. Her right hand keeps hold
of the rein at the red saddle the whole time. Nothing about her drifts,
morphs or re-rolls.
The camera is her own hand at the end of her outstretched arm, so it rides
with her: a light regular bounce in time with the horse's stride, a small
amount of hand sway, nothing more. It never pans, tilts, zooms, pushes in or
pulls back, and it never leaves her hand. Her arm's length never changes, so
she is the same size in the frame in the last second as in the first,
filling the frame from the waist up and slightly off centre.
Show that the horse is moving through the horizon rather than through its
legs: the black mane in the lower right lifts and falls in time with the
bounce, her hair streams back off her shoulders, the tassels at her belt
swing, and the grassland, the fence and the two white gers behind her drift
steadily past. Keep the horse's legs and head out of the frame exactly as
they are now.
0.0s — She holds the frame, riding, smiling into the lens as her hair blows
across her face and she lets it.
3.0s — She laughs — a real open laugh with her eyes creasing and her
shoulders shaking, chin lifting a little, delighted. She keeps her eyes on
the lens and keeps hold of the rein.
6.5s — Still laughing, she looks off past the camera toward the hills for a
moment, then back to the lens, and the laugh eases down into a wide open
grin as she rides on.
Land: rolling green-brown steppe under a low overcast sky, layered grey
cloud, dark hills along the horizon, two white gers and a wooden fence on
the ridge behind her, flat soft daylight with no sun on her face. The
weather does not change.
Sound: her laughing, wind across an open plain, the muffled rhythm of hooves
on grass, and the creak of a saddle. No music, no voiceover, no other
people, no speech — she laughs but she never says a word and never mouths
one.
Style: photorealistic phone selfie footage shot at arm's length from
horseback, real handheld bounce, real blinking, breath in her shoulders,
hair with weight to it, fine sensor grain, 24fps.
Avoid: cut, hard cut, jump cut, dissolve, cross-fade, camera move, zoom,
pan, push in, pull back, a change of shot size, the camera leaving her hand,
a drone shot, a wide landscape shot, the horse's legs, the horse's head, a
horse rearing or bolting, a second rider, a second horse, another person, a
hat, dropping the rein, a moving mouth forming words, text overlay,
subtitles, captions, watermark, logo, timestamp, readable writing anywhere
in frame.Nine shots and not one cut
Pipeline A at full length, and the clearest example on this page of what the model will not do.
- Length
- 15s, one take
- Frame
- 9:16, 720p
- Supplied
- One drawn first frame
- Cost
- 4 + 100
Frame one has to hold a face, so the film opens on a sprint at the lens.
Fifteen seconds, nine shots, a hundred credits — and zero cuts at every detection threshold from 0.3 down to 0.1. Not fewer than asked: none. All nine shot changes are cross-fades, each about a third of a second long.
The beats are all there and all in order: the server aisle, the deck close-up,
the stairwell, the puddle, the searchlit alley, the shutter, the rooftop leap,
the ledge catch, the kneel. The brief asked for hard cuts eight times —
three HARD CUT.s inside the beats, once in the style line, once in Avoid:.
The same brief on Seedance cut eight times. This is one clip showing both the
negation rule and the cut rule.
Two more findings sit in it. The prop drifted while the pixels held: frame one has a flat matte-black deck with a hot-pink X on the lid, and by the second second she is carrying a red-cased smartphone. And the ending slows down — the closing beat asked her to pocket the deck and push up out of the crouch to run again, and the last 1.4 seconds run 14.5 → 4.4 on frame-to-frame difference. It did the first half.
The one thing that held completely is what the run was built to test. Fifteen seconds, nine shots, no reference channel at all, and her face, the ponytail, the white X on the hoodie, the hip chain and the pink-soled high-tops are all still recognisable at fourteen seconds. A first frame did a reference channel’s job for a quarter of a minute.
The brief, in full
Reproduced exactly as it was submitted, including the eight cut instructions that did nothing. Six timestamped windows carry nine shots, and all nine shot changes came out as cross-fades.
Style: photorealistic cinematic night action, handheld and urgent, in a
rain-soaked neon city. Hard fast cutting — NINE shots in fifteen seconds,
every cut a hard cut, no dissolves or fades. Deep blacks, wet reflective
concrete, hot magenta and cyan sign light, cold blue shadow, heavy rain,
steam, sparks. Real motion blur on the fast whips, restrained grain.
Vertical 9:16.
Subject: JANA — the young woman in the opening frame, identical in every
shot: blonde hair with hot-pink streaks in a high messy ponytail, pale skin,
grey-violet eyes, a small inked black X with three dots under her left eye,
a black studded choker, an oversized black hoodie with a white brushed X on
the chest, black cargo trousers with grey straps and a hanging chain, chunky
black-white-and-pink high-top sneakers. Her face, hair, marking and outfit
never drift, morph or re-roll. She is moving hard in every shot and never
stands still. She never speaks.
Prop: THE DECK — a flat matte-black handheld computer with a hot-pink X on
the lid and a glowing pink edge strip. It is in her hand or on a strap
across her body in every shot.
Action Sequence — nine shots, hard cuts:
0.0s - 2.5s Dark server aisle, red alarm strobes. Camera racing
backwards ahead of her: JANA sprints straight at the lens in the black
hoodie, deck in one hand, ponytail thrown up by the movement. HARD CUT.
MACRO low on her hands: her thumb drags across the deck's glowing pink
screen as she runs, a red strobe snapping across her knuckles.
2.5s - 5.0s Concrete stairwell, camera at the bottom of the well
looking straight up. JANA vaults the handrail two floors up, drops a full
flight through frame with the deck clamped to her chest, and slams onto
the landing in a crouch, sneakers skidding.
5.0s - 7.5s Alley, hammering rain. GROUND-LEVEL MACRO on her
black-white-and-pink sneakers smashing through a magenta-lit puddle. HARD
CUT. Camera racing backwards ahead of her: she sprints into the lens, hood
up, breath fogging, searchlights swinging over wet brick behind her.
7.5s - 10.0s Loading dock, camera flat on the wet ground. A steel roller
shutter drops fast; JANA slides under it on her hip with the deck hugged
to her chest, and it crashes shut behind her sneakers.
10.0s - 12.5s Rooftop, city glow below. Tracking side-on: she runs the
wet parapet in the black hoodie and launches over the gap between two
buildings, arms wide, ponytail streaming. HARD CUT. TIGHT on her hand
slapping the far ledge and gripping, her body slamming into wet brick.
12.5s - 15.0s Lower roof, rain through neon, camera low and close. She
hauls herself over the ledge onto one knee, flips the deck open and
hot-pink screen light hits her face and the inked X under her eye. She
looks up past the lens, grins hard, shoves the deck into her hoodie pocket
and pushes up out of the crouch to run again.
Audio: an alarm klaxon and server fans; her hard breathing throughout;
sneakers on concrete, steel and wet asphalt; rain hammering metal; a steel
shutter crashing down; wind at the roof edge; a body hitting brick; the deck
snapping open. No music, no score, no narration, no dialogue, nobody
speaking.
Avoid: no on-screen text, subtitles, captions, watermark, logo, HUD or
interface overlay; no character sheet, model sheet, turnaround, multiple
panels, side-by-side figures, split screen, white studio background,
watercolour splash or colour swatches; no standing still, posing or static
shot; no slow motion, speed ramp or freeze frame; no dissolve, fade or wipe;
no guns, gunfire, blood or injury; no second woman resembling her; no change
of hair colour, hairstyle or clothing; no visible camera or crew; no
morphing or drifting of her face, eye marking or clothes.Everything here runs in the studio.
Every model named on this page is one you can pick from the composer. New accounts start with credits to spend on exactly this.
Start picturing