20 min readMarkdown
Two pictures, fifteen seconds: the first-and-last-frame method
Make the opening frame and the closing frame yourself, hand a video model both, and let it fill the middle. Nine models take a last frame. What it buys.
Most video models take a picture and start from it. Several will also take a second picture and end on it. That second slot is the least-used control in the catalog, and it changes the job: instead of describing an ending and hoping, you make the ending, look at it, fix it for three credits, and only then pay for motion.
Here is everything we handed over for the fifteen-second clip further down. Left is second zero. Right is second fifteen. There was no third input.
Both made on Seedream 5 Lite at three credits each. Everything between them is one Seedance 2.5 take.
What the second picture buys
The ending stops being a gamble. A brief’s last beat is the one the model has the least budget left for and the most freedom over. Supplied as a picture it is not a request, it is a fixed point the take has to arrive at. Ours is a close-up of a filthy, unimpressed bird, and that is the frame the clip lands on.
The camera move stops needing to be named. Frame one is a wide shot and frame two is a medium close-up of the same set on the same axis. Nobody has to write “dolly in” and hope the model shares your definition of it — the move is the only way to get from one picture to the other, so it gets derived rather than interpreted.
Identity cannot drift, because both ends are pinned. The usual failure of a character clip is a face that is subtly not the same face by the end. With the last frame supplied, drift has nowhere to accumulate to.
And the expensive mistakes get cheap. Video is priced by the second. Three credits is what it costs to find out that your closing expression reads as sad rather than disgusted. Two hundred and eighty-nine is what it costs to find out the same thing in a finished take.
The frames are the film. The prompt is stage direction between them.
Which models take a last frame
Nine of the catalog’s video models accept one, and the ceiling on a single take runs from eight seconds to thirty across them. Costs below are for one take at that model’s longest duration.
| Model | Longest take | Last frame | Cost at its longest |
|---|---|---|---|
| Seedance 2.5 | 30s | yes | 578 cr · 720p |
| FLUX 3 | 20s | yes | 283.5 cr · 720p |
| Seedance 2.0 | 15s | yes | 225 cr · 720p |
| Seedance 2.0 Fast | 15s | yes | 187.5 cr · 720p |
| Seedance 2.0 Mini | 15s | yes | 112.5 cr · 720p |
| Kling v3 Omni | 15s | yes | 210 cr · 720p, no audio |
| P-Video | 10s | yes | 17 cr · 720p |
| Veo 3.1 | 8s | yes | 133.5 cr · no audio |
| Veo 3.1 Lite | 8s | yes | 33.5 cr · 720p |
| Grok Imagine Video 1.5 | 15s | no | first frame only |
Two things fall out of that table.
P-Video at 17 credits is where you should be testing. It takes both frames, it runs to ten seconds, and it costs a seventeenth of the take we shipped. If your two frames are far enough apart that the interpolation might tear, a P-Video pass tells you so for the price of six stills.
Duration picks the model, not quality. We needed fifteen seconds. Veo 3.1 stops at eight, so the whole Veo family was out before any comparison happened. Seedance 2.0 stops at exactly fifteen, which is a ceiling you would be sitting on rather than under. Seedance 2.5 runs to thirty and makes its own sound in the same pass, so fifteen seconds is comfortably inside it. That was the entire decision.
How we made a fifteen-second one
Four steps. The first three are pictures at three credits a generation and the fourth is the only expensive one. If the frame you need is already inside a clip you have, extract it rather than generating it — the method does not care where a frame came from.
Step one: cast the character before you build the shot
If a person or an animal has to be the same in both frames, make a plain reference picture of them first — on their own, doing nothing, against a wall. Ours is a baby kagu, the flightless grey bird that is New Caledonia’s national emblem. It appears nowhere in the finished film. Its job is to be passed as a reference into the two frames that do.
The reference, then the same bird in each frame. Three credits spent once, so the crest, the beak, the feet and the proportions are decided once instead of twice.
Two habits are worth taking from that picture. Say what a reference is for and what it is not — “keep the character design, same crest, same beak and feet, same proportions, but full body, small in frame, standing” is a different instruction from handing over a file and hoping. And describe a set as a list of materials, not as a room: name the concrete, the wall and the roof rather than the building, because a room brings its contents with it and then you are arguing with a background you invited.
If you are making more than a handful of shots with the same character, stop passing a file by hand and make it a Subject. A named identity carries its picture and its description together, and cannot be the wrong file.
Step two: the opening frame
Seedream 5 Lite, 9:16 at 2K, with the character reference attached. Three credits a pass, and it took two.
The whole frame, and the half of it the joke depends on. Four men, drinks in hand, not one of them looking at the animal on its back in the mud.
The opening frame's prompt, in full
3D animated feature film style, vertical composition, wide establishing shot.
A tropical agricultural fair livestock pen: packed red dirt, weathered wooden
rail fences, and a wide muddy puddle in the middle ground.
Center of frame: an ENORMOUS Brahman bull lying on its back in the mud, all
four legs up in the air, blissfully wallowing, tongue lolling out, eyes
half-closed in pure contentment. He is obviously having the best moment of his
day. Mud splattered across his hide.
At the near edge of the puddle, tiny by comparison and no taller than the
bull's hoof: a baby kagu chick — ruffled grey crest swept back, bright orange
beak, bright orange feet, huge glossy black eyes, fluffy pale grey and cream
plumage. He is frozen mid-step, little wings half-raised in alarm, beak open,
staring at the bull in absolute horror. He believes the bull is drowning.
Background: three or four cattle farmers in classic cowboy hats — curved
upturned brim, creased pinched crown — walking past behind the fence, chatting
to each other, drinks in hand, completely ignoring the bull, not one of them
looking at the puddle. Their indifference must be clearly readable.
Bright harsh midday tropical sun, dust hanging in the air, shallow depth of
field on the background figures.
Keep the kagu chick's character design exactly as in the reference image —
same crest, same orange beak and feet, same eyes, same proportions — but full
body, small in frame, standing.
No readable text anywhere, no signage, no lettering on banners, no logos, no
subtitles, no watermark.The prompt is a wide establishing shot and it is written as a list of what is in the frame, in depth order: ground, then the bull in the middle, then the chick at the near edge, then the men behind the fence. Nothing in it is a camera term.
The onlookers are load-bearing and they are written that way.
completely ignoring the bull, not one of them looking at the puddle. Their indifference must be clearly readable is not set dressing — it is what tells
the viewer the animal is fine, which is the only reason the chick’s panic is
funny rather than alarming. Anything a scene depends on has to be described as
a fact about the picture, not implied by the situation.
The second pass was a refinement, not a second prompt. The first came back correct except for the hats, which were generic wide-brimmed rather than cowboy hats. Re-running the prompt with better wording would have re-rolled the composition with it — new bull, new puddle, a good frame thrown away to fix a detail. Refining the existing generation with one instruction that changes one element keeps everything else.
That frame was then touched up by hand, which is why re-running this prompt gives you a cousin rather than a twin. Once a frame is finished, the file is the authority and the prompt is only the recipe. Animate from the file.
Step three: the closing frame, made from the opening one
Frame one is the first reference for frame two. Everything that must not change — the fence layout, the palm at the left, the marquee at the upper right, the shadow across the dirt, the height of the sun — comes from the picture rather than from your ability to describe it twice the same way. The character reference goes in second, to hold the beak and crest under a coat of mud.
The whole frame, and the read it has to survive at phone size. The mud goes everywhere except the two features carrying the expression.
The closing frame's prompt, in full
Same scene, same camera position, same lighting — the final frame of the shot.
3D animated feature film style, vertical composition, bright harsh midday
tropical sun, red dirt, wooden rail fences, same agricultural fair pen.
THE CHICK: the baby kagu is now COMPLETELY caked in mud — his pale grey down
plastered brown and dripping, his swept-back crest flattened and gloopy, thick
mud sliding off his wingtips, orange feet buried in slop. He stands square to
camera in the foreground right.
His face is the subject of the frame — a large, clear, unobstructed close read
of his expression: DISGUSTED. Deadpan, flat, thoroughly unimpressed. Eyes
half-lidded and looking directly into the lens, beak closed and slightly
downturned in a small grimace, head tilted a fraction to one side, one wing
lifted away from his body as if he cannot bear to touch himself. A single
strand of mud drips slowly off the tip of his crest. He is not sad, not
defeated, not crying — he is quietly revolted, and he blames the camera. Dry
comedic deadpan, the look-to-camera beat.
Despite the mud his huge glossy dark eyes and orange beak stay clean and
perfectly legible — the expression must read instantly at phone size. Framing
is noticeably tighter on him than the opening wide shot: he fills more of the
lower frame, chest-up to full body, face clearly the focal point, sharp focus.
THE BULL: gone. Only his muddy hindquarters and tail are still visible at the
very left edge of frame as he ambles calmly away, unhurried, unharmed. The mud
puddle behind the chick is churned up and empty, with a wide slick drag-trail
and deep hoofprints leading out of it toward the left.
THE COWBOYS: the same four men in cowboy hats behind the wooden fence, softly
out of focus now, still talking to each other, still holding their drink cans,
still laughing at something off-screen. Not one of them is looking at the bull,
the puddle or the chick.
Same background: blue sky with soft white clouds, palm tree at left, white
marquee tent at upper right, same fence layout, same tree shadow across the
red dirt foreground.
Keep the kagu chick's character design identical to the reference — same crest
shape, same orange beak, same orange feet, same eye size, same proportions.
Only the mud and the expression are new.
No readable text anywhere, no signage lettering, no logos, no subtitles,
no watermark.Three decisions inside that prompt do most of the work.
Frame the ending tighter than the opening. The difference between the two crops is the camera move. A last frame at the same size as the first leaves the model nothing to do but wait.
Say what the expression is, and then say what it is not. “Disgusted, deadpan, half-lidded, beak downturned — not sad, not defeated, not crying” is the line that separates a joke from a tragedy. Disgust also reads in a single frame with no story attached, where triumph needs the viewer to infer what the character believes. This is the frame a thumbnail gets cut from, so it has to work with the sound off and the context missing.
Protect the face from the thing you just asked for. Mud that covers the
character covers the performance with it. his huge glossy dark eyes and orange beak stay clean and perfectly legible is what keeps the punchline readable at
phone size, and it is the correction this frame needed most.
Step four: the clip
Both frames go in, the duration goes in, and the prompt describes only what happens between them.
- Length
- 15s, one take
- Frame
- 9:16, 720p
- Supplied
- Two frames
- Cost
- 289
The two frames at the top of this page, and nothing else. The sound was made in the same run.
The clip's brief, in full
Animate between the supplied first and last frames. 3D animated feature film
look, vertical 9:16, 15 seconds, one continuous take — no cuts, no shot changes.
CAMERA — starts locked on the wide establishing frame with faint handheld
drift. From about 0:10 it begins a slow, smooth dolly push-in toward the mud-
caked chick, tightening steadily and settling exactly on the final frame by
0:14: medium close-up, chick large in frame, background thrown soft. The push
is continuous and unhurried. Never cut, never zoom out, never whip-pan.
0:00–0:03 The bull keeps wallowing on his back in the puddle, legs paddling
lazily in the air, tongue lolling, eyes half-closed in bliss, mud
slopping and bubbling around him. The clean little grey kagu chick
snaps his head left, sees the bull, and freezes — beak falling open,
tiny wings flaring out in alarm.
0:03–0:06 Panic. He scurries back and forth along the near edge of the puddle,
flapping his useless wings, barking sharply and fast, glancing from
the bull to the camera and back.
0:06–0:11 The rescue. He charges into the mud, shoulders into the bull's flank
and shoves with his whole body — orange feet skidding backwards
through the slop, splattering himself. He abandons pushing, clamps a
fold of the bull's hide in his beak and hauls, leaning right back,
straining, feet sliding. The bull slowly blinks, entirely unbothered,
and flops his legs again — a heavy wall of mud slaps the chick full
in the face. His barking stops dead. He is now filthy, grey down
plastered brown, a thick glob of mud sliding down his crest.
0:11–0:14 The bull rolls onto his front, stands up entirely on his own, shakes
his hide, and ambles calmly out of frame to the LEFT — only his muddy
hindquarters and swinging tail still visible at the edge. The chick
trudges back out of the puddle and watches him go. The camera pushes
in on him as he does.
0:14–0:15 He turns square to camera and holds a flat, deadpan, thoroughly
DISGUSTED stare straight down the lens — eyes half-lidded, beak shut
and downturned in a small grimace, one wing held stiffly away from
his own filthy body as if he cannot bear to touch himself. A strand
of mud drips slowly off the tip of his crest. Not sad, not defeated,
not crying — quietly revolted, and he blames the camera. Hold on this.
THROUGHOUT — the four cowboys behind the wooden fence never once look at the
bull, the puddle or the chick. They keep talking to each other, sipping their
cans, laughing at something off-screen, for the entire fifteen seconds. Their
complete indifference must stay readable even as they drift out of focus during
the push-in. Nobody enters the pen. Nobody intervenes. Nobody is worried.
THE BULL IS NEVER IN DISTRESS at any point — he is blissfully enjoying his mud
bath from the first frame to the moment he strolls away. No struggling, no
panic, no sinking, no rescue actually occurring.
Character consistency: keep the kagu chick's design exactly as supplied —
swept-back grey crest, orange beak, orange feet, huge glossy dark eyes, same
proportions. Only the mud accumulates, progressively, across the shot.
Audio: the chick's sharp dog-like barking, frantic then cutting abruptly to
silence the instant the mud hits him; the bull's low contented grunts and
snorts; thick wet mud squelching, sucking and splashing; one heavy wet plop as
mud slides off the crest at the end; warm outdoor fair ambience; a distant
rodeo PA announcer and country music; sparse far-off applause and laughter; the
cowboys murmuring and chuckling throughout. No music score, no voice-over.
No on-screen text, no captions, no subtitles, no signage lettering, no logos,
no watermark.The brief is written as a timeline because both ends are already fixed: every line in it is about the middle. Note what it does not contain — no description of the set, no colours for the character, no framing for the final beat. All of that is in the pictures, and repeating it in words hands the model a second, differently-worded opinion to reconcile with the first.
That still comes to 3,917 characters against the 4,000 Seedance 2.5 accepts. A timeline is long by nature, so the ceiling is a real constraint on this method rather than a theoretical one — which is another reason not to spend any of it re-describing a picture you have already supplied.
What actually came back
We measured the delivered file against the brief. Four instructions landed exactly. Four did not, and they are the more useful half.
One continuous take, no cut, landing on frame two. The push-in is smooth and unbroken and finishes on the supplied close-up. The single biggest risk of this method — a visible tear where the model gives up and jumps to the destination — did not happen.
Not one of the four cowboys ever looks. Fifteen seconds of complete indifference, which is the instruction the joke rests on. They are also barely animated: near-static idle poses behind the fence. The refusal held; the performance did not.
The barking stops dead on the splash. The one audio beat with a visible cause in the picture arrived on the frame it was asked for.
No readable text anywhere. No signage, no lettering, no watermark, in any frame.
And the four that missed:
The push-in started at 0:00, not 0:10. The brief asked for a locked frame that begins moving two-thirds of the way through. The camera moved steadily from the first picture to the last for the whole fifteen seconds instead. A timing instruction inside a shot is a suggestion; the shape of the move is the instruction that lands. If a move has to start late, it needs its own generation.
A two-part action became one part. The rescue was written as shoulder into the flank, fail, then haul on the hide with the beak. Only the haul survived. Two consecutive physical actions in one beat is a choice you are handing to the model, and it will make it.
The mud arrived in a single hit at 0:08, not progressively. Only the mud accumulates, progressively, across the shot is the hardest thing to buy with
this method, and the reason is structural: the two frames define a clean state
and a filthy state, and nothing obliges the model to spend more than one moment
moving between them. A gradual change across a take wants an intermediate
frame, not an adverb.
The PA announcer said words. Asking for a distant rodeo PA announcer
produced synthesised English — an intelligible, meaningless sentence under the
last five seconds. Ambience described as a person talking gets you a person
talking, and invented speech is the loudest artificial tell a clip can carry.
Ask for crowd noise with no announcer in it, or write the line yourself.
The mix is thinner than the brief in general: the country music, the applause and the laughter never arrived, and the track measures −26.0 LUFS integrated with a −5.5 dBFS true peak — sparse and quiet rather than a fairground. Native audio is worth having and is not yet worth handing a sound design to.
What breaks, and the fix
| Symptom | Cause | Fix |
|---|---|---|
| The take tears or jumps near the end | The two frames are too far apart to interpolate | Check both share the set and the camera axis. If the crop gap is large, make a middle frame, run two takes and join them |
| The character drifts between the frames | Identity was described rather than referenced | Pass the same character reference into both frames, and say what the reference is for |
| The closing expression is lost under an effect | Mud, rain, smoke or shadow applied to the whole face | Ask for the effect and exempt the eyes and mouth by name |
| A background element contradicts the setting | The set was named as a place instead of as materials | Describe the surfaces. A place brings its contents with it |
| Words appear on a sign or a banner | The standard failure of any model shown a public space | Put the negatives at the head of the prompt rather than the tail, and re-run |
| A sound runs past the beat it should stop on | An audio cue with no visible cause in the picture | Fix it in an edit. Cut the track — do not spend a second take on it |
| The camera does something nobody asked for | The motion was named instead of shown | The frames are the move. Change the crops, not the adjectives |
What it cost
| Credits | |
|---|---|
| Character reference · Seedream 5 Lite | 3 |
| Opening frame, two passes | 6 |
| Closing frame | 3 |
| The clip · Seedance 2.5, 15s at 720p | 289 |
| Total | 301 |
Four per cent of that went on the pictures and ninety-six on the motion, and that ratio is the argument for the whole method: every decision that could be made inside the four per cent was made there. Price the take before you run it — a fifteen-second Seedance 2.5 clip is exactly three times a five-second one, and nothing on the duration slider says so.
Everything here runs in the studio.
Every model named on this page is one you can pick from the composer. New accounts start with credits to spend on exactly this.
Start picturing