---
title: "Two pictures, fifteen seconds: the first-and-last-frame method"
description: "Make the opening frame and the closing frame yourself, hand a video model both, and let it fill the middle. Nine models take a last frame. What it buys."
canonical_url: "https://justpictur.ing/guides/first-and-last-frame/"
published: 2026-08-17
updated: 2026-08-17
locale: "en"
models: ["seedance-2.5", "seedance-2.0", "seedance-2.0-fast", "seedance-2.0-mini", "veo-3.1", "veo-3.1-lite", "kling-v3-omni-video", "flux-3", "p-video", "grok-imagine-video-1.5", "seedream-5-lite"]
---

Most video models take a picture and start from it. Several will also take a
second picture and **end** on it. That second slot is the least-used control in
the catalog, and it changes the job: instead of describing an ending and hoping,
you make the ending, look at it, fix it for three credits, and only then pay for
motion.

Here is everything we handed over for the fifteen-second clip further down. Left
is second zero. Right is second fifteen. There was no third input.

<figure class="guide-example guide-example--wide">
<img class="guide-example-plate" src="/showcase/guides/the-rescue-two-frames.webp" width="1104" height="960" loading="lazy" decoding="async" alt="Two vertical frames side by side. Left: a bull on its back in a mud puddle with all four legs up, a small clean grey chick alarmed at the near edge, four cowboys chatting behind a fence. Right: the same chick in close-up, caked in mud, glaring at the camera while the bull walks out of frame at the left." />
<figcaption class="guide-example-side"><p class="guide-example-input-note">Both made on Seedream 5 Lite at three credits each. Everything between them is one Seedance 2.5 take.</p></figcaption>
</figure>

## What the second picture buys

**The ending stops being a gamble.** A brief's last beat is the one the model
has the least budget left for and the most freedom over. Supplied as a picture
it is not a request, it is a fixed point the take has to arrive at. Ours is a
close-up of a filthy, unimpressed bird, and that is the frame the clip lands on.

**The camera move stops needing to be named.** Frame one is a wide shot and
frame two is a medium close-up of the same set on the same axis. Nobody has to
write "dolly in" and hope the model shares your definition of it — the move is
the only way to get from one picture to the other, so it gets derived rather
than interpreted.

**Identity cannot drift, because both ends are pinned.** The usual failure of a
character clip is a face that is subtly not the same face by the end. With the
last frame supplied, drift has nowhere to accumulate to.

**And the expensive mistakes get cheap.** Video is priced by the second. Three
credits is what it costs to find out that your closing expression reads as sad
rather than disgusted. Two hundred and eighty-nine is what it costs to find out
the same thing in a finished take.

> The frames are the film. The prompt is stage direction between them.

## Which models take a last frame

Nine of the catalog's video models accept one, and the ceiling on a single take
runs from eight seconds to thirty across them. Costs below are for one take at
that model's longest duration.

| Model                                                     | Longest take | Last frame | Cost at its longest     |
| --------------------------------------------------------- | ------------ | ---------- | ----------------------- |
| [Seedance 2.5](/models/seedance-2.5/)                     | 30s          | yes        | 578 cr · 720p           |
| [FLUX 3](/models/flux-3/)                                 | 20s          | yes        | 283.5 cr · 720p         |
| [Seedance 2.0](/models/seedance-2.0/)                     | 15s          | yes        | 225 cr · 720p           |
| [Seedance 2.0 Fast](/models/seedance-2.0-fast/)           | 15s          | yes        | 187.5 cr · 720p         |
| [Seedance 2.0 Mini](/models/seedance-2.0-mini/)           | 15s          | yes        | 112.5 cr · 720p         |
| [Kling v3 Omni](/models/kling-v3-omni-video/)             | 15s          | yes        | 210 cr · 720p, no audio |
| [P-Video](/models/p-video/)                               | 10s          | yes        | 17 cr · 720p            |
| [Veo 3.1](/models/veo-3.1/)                               | 8s           | yes        | 133.5 cr · no audio     |
| [Veo 3.1 Lite](/models/veo-3.1-lite/)                     | 8s           | yes        | 33.5 cr · 720p          |
| [Grok Imagine Video 1.5](/models/grok-imagine-video-1.5/) | 15s          | **no**     | first frame only        |

Two things fall out of that table.

**P-Video at 17 credits is where you should be testing.** It takes both frames,
it runs to ten seconds, and it costs a seventeenth of the take we shipped. If
your two frames are far enough apart that the interpolation might tear, a
P-Video pass tells you so for the price of six stills.

**Duration picks the model, not quality.** We needed fifteen seconds.
[Veo 3.1](/guides/veo-3-1-prompting/) stops at eight, so the whole Veo family
was out before any comparison happened. Seedance 2.0 stops at exactly fifteen,
which is a ceiling you would be sitting on rather than under.
[Seedance 2.5](/guides/seedance-2-5-prompting/) runs to thirty and makes its own
sound in the same pass, so fifteen seconds is comfortably inside it. That was
the entire decision.

## How we made a fifteen-second one

Four steps. The first three are pictures at three credits a generation and the
fourth is the only expensive one. If the frame you need is
already inside a clip you have,
[extract it](/guides/extract-a-frame-from-a-video/) rather than generating it —
the method does not care where a frame came from.

### Step one: cast the character before you build the shot

If a person or an animal has to be the same in both frames, make a plain
reference picture of them first — on their own, doing nothing, against a wall.
Ours is a baby kagu, the flightless grey bird that is New Caledonia's national
emblem. It appears nowhere in the finished film. Its job is to be passed as a reference
into the two frames that do.

<figure class="guide-example guide-example--wide">
<img class="guide-example-plate" src="/showcase/guides/the-rescue-identity.webp" width="1488" height="480" loading="lazy" decoding="async" alt="Three crops of the same baby kagu chick: the plain reference picture, the alarmed clean chick from the opening frame, and the mud-caked glare from the closing frame. The crest, beak and eyes are identical in all three." />
<figcaption class="guide-example-side"><p class="guide-example-input-note">The reference, then the same bird in each frame. Three credits spent once, so the crest, the beak, the feet and the proportions are decided once instead of twice.</p></figcaption>
</figure>

Two habits are worth taking from that picture. **Say what a reference is for and
what it is not** — "keep the character design, same crest, same beak and feet,
same proportions, but full body, small in frame, standing" is a different
instruction from handing over a file and hoping. And **describe a set as a list
of materials, not as a room**: name the concrete, the wall and the roof rather
than the building, because a room brings its contents with it and then you are
arguing with a background you invited.

If you are making more than a handful of shots with the same character, stop
passing a file by hand and make it a Subject. A named identity carries its
picture and its description together, and cannot be the wrong file.

### Step two: the opening frame

Seedream 5 Lite, 9:16 at 2K, with the character reference attached. Three
credits a pass, and it took two.

<figure class="guide-example guide-example--wide">
<img class="guide-example-plate" src="/showcase/guides/the-rescue-opening.webp" width="1048" height="540" loading="lazy" decoding="async" alt="The opening frame beside a magnified crop of the four men behind the fence: two facing each other in conversation, one turned away, one holding up a drink, none of them looking down at the bull in the mud below them." />
<figcaption class="guide-example-side"><p class="guide-example-input-note">The whole frame, and the half of it the joke depends on. Four men, drinks in hand, not one of them looking at the animal on its back in the mud.</p></figcaption>
</figure>

<details class="guide-fold"><summary>The opening frame's prompt, in full</summary><pre><code>
3D animated feature film style, vertical composition, wide establishing shot.&#10;&#10;A tropical agricultural fair livestock pen: packed red dirt, weathered wooden&#10;rail fences, and a wide muddy puddle in the middle ground.&#10;&#10;Center of frame: an ENORMOUS Brahman bull lying on its back in the mud, all&#10;four legs up in the air, blissfully wallowing, tongue lolling out, eyes&#10;half-closed in pure contentment. He is obviously having the best moment of his&#10;day. Mud splattered across his hide.&#10;&#10;At the near edge of the puddle, tiny by comparison and no taller than the&#10;bull's hoof: a baby kagu chick — ruffled grey crest swept back, bright orange&#10;beak, bright orange feet, huge glossy black eyes, fluffy pale grey and cream&#10;plumage. He is frozen mid-step, little wings half-raised in alarm, beak open,&#10;staring at the bull in absolute horror. He believes the bull is drowning.&#10;&#10;Background: three or four cattle farmers in classic cowboy hats — curved&#10;upturned brim, creased pinched crown — walking past behind the fence, chatting&#10;to each other, drinks in hand, completely ignoring the bull, not one of them&#10;looking at the puddle. Their indifference must be clearly readable.&#10;&#10;Bright harsh midday tropical sun, dust hanging in the air, shallow depth of&#10;field on the background figures.&#10;&#10;Keep the kagu chick's character design exactly as in the reference image —&#10;same crest, same orange beak and feet, same eyes, same proportions — but full&#10;body, small in frame, standing.&#10;&#10;No readable text anywhere, no signage, no lettering on banners, no logos, no&#10;subtitles, no watermark.</code></pre></details>

The prompt is a wide establishing shot and it is written as a list of what is in
the frame, in depth order: ground, then the bull in the middle, then the chick
at the near edge, then the men behind the fence. Nothing in it is a camera term.

**The onlookers are load-bearing and they are written that way.**
`completely ignoring the bull, not one of them looking at the puddle. Their
indifference must be clearly readable` is not set dressing — it is what tells
the viewer the animal is fine, which is the only reason the chick's panic is
funny rather than alarming. Anything a scene depends on has to be described as
a fact about the picture, not implied by the situation.

**The second pass was a refinement, not a second prompt.** The first came back
correct except for the hats, which were generic wide-brimmed rather than cowboy
hats. Re-running the prompt with better wording would have re-rolled the
composition with it — new bull, new puddle, a good frame thrown away to fix a
detail. Refining the existing generation with one instruction that changes one
element keeps everything else.

<blockquote class="guide-quote--caution">
That frame was then touched up by hand, which is why re-running this prompt
gives you a cousin rather than a twin. Once a frame is finished,
<strong>the file is the authority and the prompt is only the recipe</strong>.
Animate from the file.
</blockquote>

### Step three: the closing frame, made from the opening one

**Frame one is the first reference for frame two.** Everything that must not
change — the fence layout, the palm at the left, the marquee at the upper right,
the shadow across the dirt, the height of the sun — comes from the picture
rather than from your ability to describe it twice the same way. The character
reference goes in second, to hold the beak and crest under a coat of mud.

<figure class="guide-example guide-example--wide">
<img class="guide-example-plate" src="/showcase/guides/the-rescue-closing.webp" width="968" height="540" loading="lazy" decoding="async" alt="The closing frame beside a magnified crop of the chick's face: mud plastered across the crest and body, while the eyes, brows and orange beak stay clean and sharply legible." />
<figcaption class="guide-example-side"><p class="guide-example-input-note">The whole frame, and the read it has to survive at phone size. The mud goes everywhere except the two features carrying the expression.</p></figcaption>
</figure>

<details class="guide-fold"><summary>The closing frame's prompt, in full</summary><pre><code>
Same scene, same camera position, same lighting — the final frame of the shot.&#10;&#10;3D animated feature film style, vertical composition, bright harsh midday&#10;tropical sun, red dirt, wooden rail fences, same agricultural fair pen.&#10;&#10;THE CHICK: the baby kagu is now COMPLETELY caked in mud — his pale grey down&#10;plastered brown and dripping, his swept-back crest flattened and gloopy, thick&#10;mud sliding off his wingtips, orange feet buried in slop. He stands square to&#10;camera in the foreground right.&#10;&#10;His face is the subject of the frame — a large, clear, unobstructed close read&#10;of his expression: DISGUSTED. Deadpan, flat, thoroughly unimpressed. Eyes&#10;half-lidded and looking directly into the lens, beak closed and slightly&#10;downturned in a small grimace, head tilted a fraction to one side, one wing&#10;lifted away from his body as if he cannot bear to touch himself. A single&#10;strand of mud drips slowly off the tip of his crest. He is not sad, not&#10;defeated, not crying — he is quietly revolted, and he blames the camera. Dry&#10;comedic deadpan, the look-to-camera beat.&#10;&#10;Despite the mud his huge glossy dark eyes and orange beak stay clean and&#10;perfectly legible — the expression must read instantly at phone size. Framing&#10;is noticeably tighter on him than the opening wide shot: he fills more of the&#10;lower frame, chest-up to full body, face clearly the focal point, sharp focus.&#10;&#10;THE BULL: gone. Only his muddy hindquarters and tail are still visible at the&#10;very left edge of frame as he ambles calmly away, unhurried, unharmed. The mud&#10;puddle behind the chick is churned up and empty, with a wide slick drag-trail&#10;and deep hoofprints leading out of it toward the left.&#10;&#10;THE COWBOYS: the same four men in cowboy hats behind the wooden fence, softly&#10;out of focus now, still talking to each other, still holding their drink cans,&#10;still laughing at something off-screen. Not one of them is looking at the bull,&#10;the puddle or the chick.&#10;&#10;Same background: blue sky with soft white clouds, palm tree at left, white&#10;marquee tent at upper right, same fence layout, same tree shadow across the&#10;red dirt foreground.&#10;&#10;Keep the kagu chick's character design identical to the reference — same crest&#10;shape, same orange beak, same orange feet, same eye size, same proportions.&#10;Only the mud and the expression are new.&#10;&#10;No readable text anywhere, no signage lettering, no logos, no subtitles,&#10;no watermark.</code></pre></details>

Three decisions inside that prompt do most of the work.

**Frame the ending tighter than the opening.** The difference between the two
crops _is_ the camera move. A last frame at the same size as the first leaves
the model nothing to do but wait.

**Say what the expression is, and then say what it is not.** "Disgusted,
deadpan, half-lidded, beak downturned — not sad, not defeated, not crying" is
the line that separates a joke from a tragedy. Disgust also reads in a single
frame with no story attached, where triumph needs the viewer to infer what the
character believes. This is the frame a thumbnail gets cut from, so it has to
work with the sound off and the context missing.

**Protect the face from the thing you just asked for.** Mud that covers the
character covers the performance with it. `his huge glossy dark eyes and orange
beak stay clean and perfectly legible` is what keeps the punchline readable at
phone size, and it is the correction this frame needed most.

### Step four: the clip

Both frames go in, the duration goes in, and the prompt describes only what
happens between them.

<figure class="guide-example">
<div class="guide-example-stage"><video src="/showcase/guides/the-rescue.mp4" poster="/showcase/guides/the-rescue.webp" width="608" height="1080" controls controlslist="nodownload noplaybackrate noremoteplayback" disablepictureinpicture playsinline preload="none" aria-label="A tiny kagu chick panics at a bull wallowing on its back in a mud puddle, hauls at the bull's hide with his beak, is hit full in the face by mud, and ends staring flatly into the lens as the bull strolls away unharmed"></video></div>
<figcaption class="guide-example-side">
<dl class="guide-example-spec">
<div><dt>Length</dt><dd>15s, one take</dd></div>
<div><dt>Frame</dt><dd>9:16, 720p</dd></div>
<div><dt>Supplied</dt><dd>Two frames</dd></div>
<div><dt>Cost</dt><dd><span class="credit">289</span></dd></div>
</dl>
<p class="guide-example-input-note">The two frames at the top of this page, and nothing else. The sound was made in the same run.</p>
</figcaption>
</figure>

<details class="guide-fold"><summary>The clip's brief, in full</summary><pre><code>
Animate between the supplied first and last frames. 3D animated feature film&#10;look, vertical 9:16, 15 seconds, one continuous take — no cuts, no shot changes.&#10;&#10;CAMERA — starts locked on the wide establishing frame with faint handheld&#10;drift. From about 0:10 it begins a slow, smooth dolly push-in toward the mud-&#10;caked chick, tightening steadily and settling exactly on the final frame by&#10;0:14: medium close-up, chick large in frame, background thrown soft. The push&#10;is continuous and unhurried. Never cut, never zoom out, never whip-pan.&#10;&#10;0:00–0:03  The bull keeps wallowing on his back in the puddle, legs paddling&#10;           lazily in the air, tongue lolling, eyes half-closed in bliss, mud&#10;           slopping and bubbling around him. The clean little grey kagu chick&#10;           snaps his head left, sees the bull, and freezes — beak falling open,&#10;           tiny wings flaring out in alarm.&#10;&#10;0:03–0:06  Panic. He scurries back and forth along the near edge of the puddle,&#10;           flapping his useless wings, barking sharply and fast, glancing from&#10;           the bull to the camera and back.&#10;&#10;0:06–0:11  The rescue. He charges into the mud, shoulders into the bull's flank&#10;           and shoves with his whole body — orange feet skidding backwards&#10;           through the slop, splattering himself. He abandons pushing, clamps a&#10;           fold of the bull's hide in his beak and hauls, leaning right back,&#10;           straining, feet sliding. The bull slowly blinks, entirely unbothered,&#10;           and flops his legs again — a heavy wall of mud slaps the chick full&#10;           in the face. His barking stops dead. He is now filthy, grey down&#10;           plastered brown, a thick glob of mud sliding down his crest.&#10;&#10;0:11–0:14  The bull rolls onto his front, stands up entirely on his own, shakes&#10;           his hide, and ambles calmly out of frame to the LEFT — only his muddy&#10;           hindquarters and swinging tail still visible at the edge. The chick&#10;           trudges back out of the puddle and watches him go. The camera pushes&#10;           in on him as he does.&#10;&#10;0:14–0:15  He turns square to camera and holds a flat, deadpan, thoroughly&#10;           DISGUSTED stare straight down the lens — eyes half-lidded, beak shut&#10;           and downturned in a small grimace, one wing held stiffly away from&#10;           his own filthy body as if he cannot bear to touch himself. A strand&#10;           of mud drips slowly off the tip of his crest. Not sad, not defeated,&#10;           not crying — quietly revolted, and he blames the camera. Hold on this.&#10;&#10;THROUGHOUT — the four cowboys behind the wooden fence never once look at the&#10;bull, the puddle or the chick. They keep talking to each other, sipping their&#10;cans, laughing at something off-screen, for the entire fifteen seconds. Their&#10;complete indifference must stay readable even as they drift out of focus during&#10;the push-in. Nobody enters the pen. Nobody intervenes. Nobody is worried.&#10;&#10;THE BULL IS NEVER IN DISTRESS at any point — he is blissfully enjoying his mud&#10;bath from the first frame to the moment he strolls away. No struggling, no&#10;panic, no sinking, no rescue actually occurring.&#10;&#10;Character consistency: keep the kagu chick's design exactly as supplied —&#10;swept-back grey crest, orange beak, orange feet, huge glossy dark eyes, same&#10;proportions. Only the mud accumulates, progressively, across the shot.&#10;&#10;Audio: the chick's sharp dog-like barking, frantic then cutting abruptly to&#10;silence the instant the mud hits him; the bull's low contented grunts and&#10;snorts; thick wet mud squelching, sucking and splashing; one heavy wet plop as&#10;mud slides off the crest at the end; warm outdoor fair ambience; a distant&#10;rodeo PA announcer and country music; sparse far-off applause and laughter; the&#10;cowboys murmuring and chuckling throughout. No music score, no voice-over.&#10;&#10;No on-screen text, no captions, no subtitles, no signage lettering, no logos,&#10;no watermark.</code></pre></details>

The brief is written as a timeline because both ends are already fixed: every
line in it is about the middle. Note what it does **not** contain — no
description of the set, no colours for the character, no framing for the final
beat. All of that is in the pictures, and repeating it in words hands the model
a second, differently-worded opinion to reconcile with the first.

That still comes to **3,917 characters against the 4,000 Seedance 2.5
[accepts](/guides/seedance-2-5-prompting/)**. A timeline is long by nature, so
the ceiling is a real constraint on this method rather than a theoretical one —
which is another reason not to spend any of it re-describing a picture you have
already supplied.

## What actually came back

We measured the delivered file against the brief. Four instructions landed
exactly. Four did not, and they are the more useful half.

**One continuous take, no cut, landing on frame two.** The push-in is smooth and
unbroken and finishes on the supplied close-up. The single biggest risk of this
method — a visible tear where the model gives up and jumps to the destination —
did not happen.

**Not one of the four cowboys ever looks.** Fifteen seconds of complete
indifference, which is the instruction the joke rests on. They are also barely
animated: near-static idle poses behind the fence. The refusal held; the
performance did not.

**The barking stops dead on the splash.** The one audio beat with a visible
cause in the picture arrived on the frame it was asked for.

**No readable text anywhere.** No signage, no lettering, no watermark, in any
frame.

And the four that missed:

**The push-in started at 0:00, not 0:10.** The brief asked for a locked frame
that begins moving two-thirds of the way through. The camera moved steadily from
the first picture to the last for the whole fifteen seconds instead. **A timing
instruction inside a shot is a suggestion; the shape of the move is the
instruction that lands.** If a move has to start late, it needs its own
generation.

**A two-part action became one part.** The rescue was written as shoulder into
the flank, fail, then haul on the hide with the beak. Only the haul survived.
**Two consecutive physical actions in one beat is a choice you are handing to
the model, and it will make it.**

**The mud arrived in a single hit at 0:08, not progressively.** `Only the mud
accumulates, progressively, across the shot` is the hardest thing to buy with
this method, and the reason is structural: the two frames define a clean state
and a filthy state, and nothing obliges the model to spend more than one moment
moving between them. **A gradual change across a take wants an intermediate
frame, not an adverb.**

**The PA announcer said words.** Asking for `a distant rodeo PA announcer`
produced synthesised English — an intelligible, meaningless sentence under the
last five seconds. Ambience described as a person talking gets you a person
talking, and invented speech is the loudest artificial tell a clip can carry.
Ask for crowd noise with no announcer in it, or write the line yourself.

The mix is thinner than the brief in general: the country music, the applause
and the laughter never arrived, and the track measures **−26.0 LUFS integrated
with a −5.5 dBFS true peak** — sparse and quiet rather than a fairground. Native
audio is worth having and is not yet worth handing a sound design to.

## What breaks, and the fix

| Symptom                                         | Cause                                                     | Fix                                                                                                                            |
| ----------------------------------------------- | --------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------ |
| The take tears or jumps near the end            | The two frames are too far apart to interpolate            | Check both share the set and the camera axis. If the crop gap is large, make a middle frame, run two takes and join them        |
| The character drifts between the frames         | Identity was described rather than referenced              | Pass the same character reference into both frames, and say what the reference is for                                           |
| The closing expression is lost under an effect  | Mud, rain, smoke or shadow applied to the whole face       | Ask for the effect and exempt the eyes and mouth by name                                                                        |
| A background element contradicts the setting    | The set was named as a place instead of as materials       | Describe the surfaces. A place brings its contents with it                                                                      |
| Words appear on a sign or a banner              | The standard failure of any model shown a public space     | Put the negatives at the head of the prompt rather than the tail, and re-run                                                    |
| A sound runs past the beat it should stop on    | An audio cue with no visible cause in the picture          | Fix it in an edit. Cut the track — do not spend a second take on it                                                             |
| The camera does something nobody asked for      | The motion was named instead of shown                      | The frames are the move. Change the crops, not the adjectives                                                                   |

## What it cost

|                                        | Credits |
| -------------------------------------- | ------- |
| Character reference · Seedream 5 Lite  | 3       |
| Opening frame, two passes              | 6       |
| Closing frame                          | 3       |
| The clip · Seedance 2.5, 15s at 720p   | 289     |
| **Total**                              | **301** |

**Four per cent of that went on the pictures and ninety-six on the motion**, and
that ratio is the argument for the whole method: every decision that could be
made inside the four per cent was made there. Price the take before you run it —
a fifteen-second Seedance 2.5 clip is exactly three times a five-second one, and
nothing on the duration slider says so.

---

## About justpictur.ing

justpictur.ing is an AI picture studio: prompt-driven images, clips, voice-overs and music, and storyboard-backed films with consistent characters across every scene.

Agents connect over MCP and can generate directly:

```
claude mcp add justpicturing https://api.justpictur.ing/mcp --transport http
```

The app is at https://app.justpictur.ing and requires an account.

Every guide on this site has a Markdown twin at its own URL plus `.md`. Every guide: https://justpictur.ing/guides/
