---
title: "Grok Imagine Video 1.5: the complete prompting guide"
description: "One picture in, up to fifteen seconds out, with sound you cannot turn off. Nine rules, three pipelines and the failures worth knowing about first."
canonical_url: "https://justpictur.ing/guides/grok-imagine-video-1-5-prompting/"
published: 2026-08-15
updated: 2026-08-16
locale: "en"
models: ["grok-imagine-video-1.5", "seedream-5-pro", "gpt-image-2"]
---

Grok Imagine Video 1.5 **animates a picture you give it**. One to fifteen
seconds, sound made in the same pass, and no way to run it from words alone:
every clip starts from an image, and that image comes back as frame one almost
exactly.

That is the whole shape of the model. There is no reference slot, no last-frame
slot and no audio switch, so most of what would be a setting on another model is
either **already in your picture** or **not available at all**.

> The picture is the brief. The prompt is only what changes.

It also builds the clip **front to back**: frame one first, then each moment
from the one before it, rather than planning the whole thing and filling it in.
Half the rules below come from that — why your opening composition survives, why
the order of your sentences is the order of the film, and why the model would
rather turn a head than cut.

## Overview

### The short version

Four decisions do most of the work. If you read nothing else, read these. Each
one is written out in full further down.

<ol class="guide-keys">
<li class="guide-key">
<p class="guide-key-title">Spend your effort on the picture, then describe only what moves</p>
<p class="guide-key-body">Your image comes back as frame one almost untouched: <strong>35.5 dB</strong> against the original, and three takes on three unrelated pictures agreed to within 0.15 dB. Framing, wardrobe, props and light all survive. Describing them again wins nothing, and invites the model to redraw something that was already right.</p>
<div class="guide-key-lines">
<div class="guide-key-quote guide-key-quote--no"><p class="guide-key-quote-label">Instead of</p><p class="guide-key-quote-text">A raccoon on a doormat at night, infrared, a white door behind it…</p></div>
<div class="guide-key-quote guide-key-quote--yes"><p class="guide-key-quote-label">Write</p><p class="guide-key-quote-text">The frame is already correct. Animate it and change nothing else.</p></div>
</div>
</li>
<li class="guide-key">
<p class="guide-key-title">Telling it what not to do does nothing</p>
<p class="guide-key-body">xAI say plainly that negative phrasing is ignored, and our runs agree. One fifteen-second brief demanded hard cuts and banned dissolves <strong>eight times over</strong>, and came back with a cross-fade at all nine of its shot changes. Ask for what you want as something that happens, never as something that must not.</p>
<div class="guide-key-lines">
<div class="guide-key-quote guide-key-quote--no"><p class="guide-key-quote-label">Instead of</p><p class="guide-key-quote-text">No camera movement, no zoom, no push-in, no drift</p></div>
<div class="guide-key-quote guide-key-quote--yes"><p class="guide-key-quote-label">Write</p><p class="guide-key-quote-text">The framing at ten seconds is identical to the framing at zero seconds — same crop, same angle, same lens</p></div>
</div>
</li>
<li class="guide-key">
<p class="guide-key-title">You cannot ask for a cut. You can make one necessary</p>
<p class="guide-key-body">Grok changes shot the way a body would. Ask for a new <em>angle</em> at the same distance and it turns a head. Ask for a new <em>distance</em> and no movement of the body can do that, so it cuts. <strong>One take with no cut instruction anywhere came back with two clean hard cuts, both of them size jumps.</strong></p>
<div class="guide-key-lines">
<p class="guide-key-practice">Want an unbroken take? Pin the shot size — <em>she is the same size in the frame in the last second as in the first</em>. Want cuts? Write shots at different distances, or render them separately and <a href="/guides/join-clips-into-one-video/">join them</a> afterwards.</p>
</div>
</li>
<li class="guide-key">
<p class="guide-key-title">The sound answers events, and there is no off switch</p>
<p class="guide-key-body">Every take makes audio whether you write for it or not. What sets the level is not your wording but <strong>how many things happen</strong>. Three takes asking only for room tone came back between <span class="term"><button class="term-word" type="button" popovertarget="def-lufs">−70 and −65 LUFS</button><span class="term-def" popover id="def-lufs">A loudness measurement. A finished social video sits around −16 LUFS; −70 LUFS is silence with a file size.</span></span> — inaudible — while a take naming a klaxon, footfalls, rain and a shutter came back at −15.8.</p>
<div class="guide-key-lines">
<p class="guide-key-practice">Write the sound as a list of events, each with a material and a surface, not as a mood. If nothing in the shot makes a noise, expect a silent file and plan to lay your own track underneath.</p>
</div>
</li>
</ol>

### What it accepts

| Input            | Limit           | Notes                                                                        |
| ---------------- | --------------- | ---------------------------------------------------------------------------- |
| Prompt           | 4,000 chars     | Roomy. Your picture is the limit here, not the character count.              |
| Starting image   | 1, **required** | It becomes frame one. There is no text-only door.                            |
| Duration         | 1–15 seconds    | Whole seconds, anywhere in the range. No automatic option.                   |
| Resolution       | 480p · 720p     | Same price, so there is no reason to shoot the smaller one.                  |
| Aspect ratio     | Auto + seven    | 16:9, 9:16, 1:1, 4:3, 3:4, 3:2, 2:3. **Auto means the picture's own shape.** |
| Reference images | **none**        | No reference channel exists. Identity comes from frame one or from nowhere.  |
| Last frame       | **none**        | You cannot pin the ending to a picture.                                      |
| Audio            | always on       | No switch, no silent mode, no separate charge.                               |

Seven named ratios is the widest set here after Seedance, and one second is a
real length rather than a rounding error — useful as a punctuating beat in an
edit. Against [Veo 3.1](/guides/veo-3-1-prompting/) and
[Seedance 2.5](/guides/seedance-2-5-prompting/) it gives up every other input:
no references, no last frame, no silence.

### What it costs

One flat rate, and the simplest price in the catalog: **about 6.7 credits a
second**, whatever the resolution and whatever the shape.

| Length | Cost                             | What it's for                                                        |
| ------ | -------------------------------- | -------------------------------------------------------------------- |
| 1s     | <span class="credit">7</span>    | A punctuating beat, or the cheapest look at how a frame moves        |
| 5s     | <span class="credit">33.5</span> | The steady length: one subject, one action, one camera idea          |
| 10s    | <span class="credit">67</span>   | A complete short piece — a vlog, a found-footage gag, a product spot |
| 15s    | <span class="credit">100</span>  | The ceiling, and the length most likely to drift                     |

<blockquote class="guide-quote--caution">
<p>480p and 720p cost the same, so <strong>a cheap test is a shorter clip, not a smaller one</strong>. On Seedance you save by dropping the resolution. Here you save by halving the length.</p>
<p>A one-second take at <span class="credit">7</span> answers a lot of questions: whether the picture is accepted, whether the ratio came out the way you set it, whether the model reads the shot the way you do.</p>
</blockquote>

---

## The rules

Nine things that are true whatever you are making. Each one came out of a take
that cost money, and they are ordered by how much they will save you.

### Describe what changes, not what the picture already shows

> The image is not a hint. It is frame one.

Measured against the still we handed in, the clip's opening frame comes back at
**35.5, 35.6 and 35.65** <span class="term"><button class="term-word" type="button" popovertarget="def-db">dB</button><span class="term-def" popover id="def-db">How close two pictures are, measured as peak signal-to-noise ratio. Above about 35 dB the difference is hard to see at all. Around 20 dB you are looking at a redrawing rather than a copy.</span></span>
across three completely different pictures — a character sprinting down a server
aisle, a bedroom selfie, a mirror shot. A model that treats a first frame as an
anchor rather than a lock scores **21.5 dB** on the same test. Fifteen decibels
apart is a different tool.

So everything the picture settles is settled: framing, wardrobe, props, the
state of the room, the light. The prompt's only job is motion.

The lock is weaker in one place, and it is worth knowing before you build a run
on it. A take whose subject wore an embroidered Mongolian deel came back at
**32.46 dB overall**, and the loss was not where you would guess:

| region                        | frame-one fidelity |
| ----------------------------- | ------------------ |
| her face                      | **34.04 dB**       |
| the grassland behind her      | 33.35 dB           |
| the deel, belt and embroidery | **27.71 dB**       |

The face is the best-preserved thing in the picture. **What gets repainted is
fabric.** The belt's medallion, its amber stone and both tassel pendants are the
same objects with the same design ten seconds later — the threads are what get
redrawn. If your subject is a garment, keep your detail at the level of the
objects on it, and do not expect 2048-pixel embroidery to arrive intact at 720p.

### Telling it what not to do does nothing

> The model ignores negative phrasing. xAI documents this, and we paid to
> confirm it.

A fifteen-second action brief wrote `HARD CUT.` three times inside its beats,
banned dissolves and fades in its style line, then banned dissolve, fade and
wipe again in its `Avoid:` line. **Eight instructions, zero hard cuts** — and a
cross-fade at all nine shot changes. The same brief on Seedance cut eight times.

The fix is never a ninth ban. It is to find the positive sentence that makes the
thing you do not want impossible:

- _"no camera movement"_ → **"the framing at ten seconds is identical to the
  framing at zero seconds — same crop, same angle, same lens, same distance from
  the mat"**;
- _"no change of shot size"_ → **"her arm's length never changes, so she is the
  same size in the frame in the last second as in the first"**;
- _"no readable text"_ → make the thing **too small and too far out of focus to
  read**, so the composition never reaches for a string in the first place.

Keep an `Avoid:` line at the foot of a brief for one reason only: a bad result
is then a result that failed **against instruction**, which is more useful to
know than one that failed in silence. Do not spend characters on it expecting
work in return.

### You cannot ask for a cut. You can make one necessary

> How a shot change is handled is decided by the content, not by the prompt.

This is the most useful thing to know about the model, and it took four takes to
pin down.

- A ten-second bedroom vlog with **no cut instruction anywhere** came back with
  two clean hard cuts, at 1.83s and 3.62s.
- A fifteen-second chase demanding nine hard cuts came back with **none**.
- Two more ten-second takes — a mirror selfie, a ride on horseback — came back
  with zero cuts at every detection threshold from 0.4 down to 0.1.

The pattern has nothing to do with the words. It is about whether a body could
have done it:

| what the next shot changes                          | what the model does                                |
| --------------------------------------------------- | -------------------------------------------------- |
| the **angle**, at the same distance                 | turns the neck, the hips, the whole body — no cut  |
| the **distance** (medium selfie → extreme close-up) | cuts, because no movement of a body moves the lens |

Both cuts in that vlog were size jumps. Both of its non-cuts were angle changes,
done as one continuous turn in front of a phone that never moved. The chase
asked for nine shot changes that were all, as the model read them, somewhere a
hand-held camera could have walked to — so it walked.

Two consequences. **For an unbroken take, pin the shot size**, and say it as
something that is true rather than something to avoid. And **a piece that really
needs eight hard cuts is eight generations**, rendered separately and
[joined afterwards](/guides/join-clips-into-one-video/), not one brief
shouting `HARD CUT.`

### The sound answers events, and there is no off switch

> `audio_mode: required`. There will be a track. The only question is whether
> anything is on it.

Seven takes, measured the same way. Sort them by what their sound block actually
named and the pattern is plain:

| what the brief's sound block asked for                 | integrated loudness |
| ------------------------------------------------------ | ------------------- |
| room tone under a silent performer                     | **−70.0 LUFS**      |
| room tone under a silent performer                     | −69.1 LUFS          |
| room tone under a silent performer                     | −64.7 LUFS          |
| night ambience plus a few small events                 | −48.7 LUFS          |
| one sustained crowd bed                                | −38.1 LUFS          |
| one vocal event — a laugh                              | −26.9 LUFS          |
| a continuous stack: klaxon, rain, footfalls, a shutter | **−15.8 LUFS**      |

The three at the top are silence with a file size. **Asked for a hum with
nothing to attach it to, the model gives you nothing at all** rather than
something quiet. Asked for a laugh, it gives a laugh at −26.9. Asked for a
chase, it gives a full mix.

So write sound the way a sound designer would: name the event, the material and
the space. _"City sounds, traffic"_ is a mood. _"Cars passing, skateboard wheels
on pavement, a distant rumble down the street"_ is a list of things that happen.

Two smaller findings worth carrying:

**A face turned away is a face with no track.** In the ride, the laugh starts at
2.86s and drops twenty decibels back to bed level at **6.40s** — the exact frame
her head turns away from the lens. She is still visibly laughing until 8.5s, so
two full seconds of it are silent. The simplest reading is that the model mixes
a voice to the **camera** rather than to the character. If a sound has to be
heard, play it to the lens.

**Ban speech twice if nobody speaks.** A smiling performer with no stated sound
is an invitation to write her a line and lip-sync it in whatever language the
model fancies. Say it in the sound block, then again as an action —
_"she never opens her mouth to speak"_. That ban held in every run here, which
makes it one of the two negatives worth writing, alongside _"no music"_.

### Ask for the end state, never an amount

> An amount reads as a gesture. A destination reads as a rule.

A broadcast recreation asked for a push-in of _"no more than three percent
across the whole clip, so slight it is nearly invisible"_ and got roughly
**twenty percent**: by the fifth second the foreground extra had left the frame.

Written as a claim about the last frame, the same instruction is exact. A
doorbell piece said: _"the framing at ten seconds is identical to the framing at
zero seconds — same crop, same angle, same lens, same distance from the mat."_
Frame 0 and frame 239 put the door edge, both hinges, the panel lines, the mat
edge, the step curve and the vignette on **identical pixels**.

This works beyond camera moves. Anything you can phrase as _"X is the same at
the end as at the beginning"_ is something the model can hold on to. Anything
phrased as an amount is a suggestion.

### A cut can land early; a body cannot

> Beats keep their order. What moves them is how long a performance takes.

Timestamps are read, and the order of your beats is the order of the film. What
varies is how far a mark slips, and that depends on what the mark _is_:

- Two shot changes that came out as **cuts** landed **early**: asked 2.0s →
  1.83s, asked 4.0s → 3.62s.
- Four that came out as **movements the subject performs** all landed **late**:
  +0.4s, +1.0s, +2.5s, +1.0s. The +2.5s pushed a whole window along with it.

A cut takes no time and can go anywhere, including ahead of its mark. A 180°
turn takes as long as a body takes, and everything queued behind it slides.

The confirmation is a take that scheduled one turn with **nothing after it**.
Its two marks landed at −0.14s and −0.10s, both early, and the turn then ran two
seconds long at no cost to the schedule, because nothing was waiting on it.

> **Budget a performed beat by how long the movement takes, not by how many
> beats fit in the runtime.** Never put a turn in the middle of a full schedule.

One more thing to expect from a beat sheet. On a format the model knows well, **a
big familiar gesture arrives whether you asked for it or not, and a small
specific one loses to it**. A brief asking for _"a quarter turn of the hips"_ got
a full turn away from the mirror, held for three seconds — the standard move of
that genre — while _"glances down at the phone for a moment"_, which is not part
of that vocabulary, never happened at all. Treat the beat sheet as a schedule the
model fills, not a script it follows.

### Spend one action on the ending

> It slows into the last second. Seedance speeds up out of it.

Over the final 1.4 seconds of the same closing beat, Grok ran **14.5 → 4.4** on
frame-to-frame difference, slowing to a stop, while Seedance ran 7.0 → 17.6 and
was at its fastest when it ended.

Neither ends on a frozen frame, but the consequence is real. That closing beat
asked for two actions: _"shoves the deck into her hoodie pocket **and** pushes up
out of the crouch to run again."_ Grok did the first and settled.

So the rule that a final block must name an event — a
[Veo](/guides/veo-3-1-prompting/) and Seedance finding — does not carry over as
written. Here the final block should name **exactly one** event, and it should
be the one you want on screen at the buzzer.

### It holds the picture, not the objects in it

> How faithful the frame is and how stable an object is are two different
> things, and only one of them is 35 dB.

The chase's first frame has the right prop: a flat matte-black handheld computer
with a hot-pink X on the lid. By the second second she is carrying a **red-cased
smartphone** with a pink X on its wallpaper, and it stays a phone through most of
the middle of the clip. A whole `Prop:` paragraph naming its shape, its colour
and its glowing edge strip did not stop that.

Text splits the same way. A broadcast frame's graphics came back with **every
string correct** — the tournament name, the score, the clock, the network
wordmark, the `LIVE` flag — and the **layout redrawn**: smaller, cell outlines
gone, a `▶` grown beside `LIVE`. That new design then held steady for the rest
of the clip. A thinner overlay in another take, a corner timestamp and a `REC`
dot, survived untouched.

Read it this way: the model copies the picture you gave it, then rebuilds every
object in it from whatever the moment before says. Anything that has to be exact
from a new angle is not safe in the prompt. **If a graphic has to be right,
composite it afterwards**, where the model never sees it.

### It will animate a photographed face

> The refusal you hit on Seedance is not here, and it is often the whole reason
> to be here.

Hand Seedance a photorealistic person and it declines: `E005 — the input or
output was flagged as sensitive`, through the reference door and the first-frame
door alike, six of six on one run and seven of seven on another. Grok takes the
same kind of picture without complaint — and since it _needs_ an image, there is
no other way to use it.

That makes it the second way around that limit, after
[Veo 3.1](/guides/veo-3-1-prompting/), and a wider one in one respect: Veo's
reference route forces 16:9, while Grok took a 4:3 frame and charged nothing
extra for the shape.

It is not a licence. A face you have no right to use is still a face you have no
right to use, and every person in the examples here was generated. The model's
own filter is also strict about **named** people, brands and logos, and it
misfires on harmless prompts. Describe a fictional person and keep brands
generic rather than arguing with it.

---

## The pipelines

Grok has **one door**: an image and a prompt. What changes between jobs is where
frame one comes from — and since it is your entire identity budget, that choice
matters more than the prompt does.

### The blocks every brief uses

Whichever way you got your picture, the prompt is built from the same parts, in
the order the model reads most reliably. It renders front to back, so **prompt
order is timeline order**: an action written at the top of the brief happens at
the top of the clip, and one buried at the bottom may arrive late or never.

```
[Continuity]  who and what must not drift, named once
[Camera]      where the lens is, and the end state that proves it stayed
[Motion]      what moves, and how hard
[Beats]       the schedule, stamped, in the order it happens
[Sound]       events, with their materials and their spaces
[Style]       the stock, the grain, the frame rate
[Avoid]       kept for the record, not for the result
```

There is **no `[Look]` block**, and that absence is the point: the light, the
colour, the wardrobe and the room all arrived with the picture. Writing them
again spends characters on a decision already made, and invites the model to
redraw it.

> **Cut anything the picture already shows.** If a sentence would be true of the
> still, it does not belong in the clip's brief.

Two more things about the beats. Keep them coarse — one subject, one action, one
camera idea each, which is xAI's own advice and matches what we measured. And
use **strong words for how hard things move**, because without them the model
picks its own reading and it is usually milder than you wanted: _"car passing"_
becomes _"car racing past at high speed"_, _"the wave crests"_ becomes _"crests
fully and pitches forward, crashing down"_.

### How to write the sound block

Sound is not optional, so this is never a block you get to skip. Two notations
work, neither better than the other, and both beat burying a sound inside a
sentence about the picture.

```
Sound: …      a paragraph among the others
AUDIO: …      a section at the very end of the prompt
```

What goes in it is **events**, each named with the material it happens on and
the space it happens in.

```
Sound: night ambience, faint crickets and a distant road. Claws on concrete
and the scuff of cardboard as the first parcel is set down. A beat of near
silence as the eyes open. Then many sets of claws on concrete at once, and
cardboard settling on cardboard as the stack builds. No music, no voices.
```

Three things worth knowing before you write one.

**Silence is not on the menu, but it is often what you get.** If nothing in the
shot makes a noise, expect a track at −65 LUFS or below and plan to score it
yourself. It also means a clip you intend to publish silent costs nothing extra
— the audio comes out of the same per-second rate either way.

**Quote a line if you want a line.** Short dialogue works, and lip-sync is
decent from a front-facing portrait with the mouth visible. Keep the line short,
keep the camera locked, and put the tone in the sound block (_"hesitant, then
determined"_).

**"No music" is one of the two negatives that seem to hold**, the other being a
ban on speech — but read that honestly. None of our takes came back scored, and
none of them was a scene with music in it. On a piece that invites a soundtrack,
expect an argument rather than an instruction.

### Pipeline A: draw frame one

The default, and what most of the examples here do. Make the still on an image
model, look at it, fix it, then hand the finished picture over.

#### What it is for

- **Best for:** anything invented — a set, a creature, a found-footage conceit,
  a piece of branded chrome, a character who exists in one file.
- **Costs you:** a second generation and a second look, usually between
  <span class="credit">4</span> and <span class="credit">11</span>.
- **Watch for:** the still is your whole identity budget, so it has to contain a
  **face** if a person is in the piece. A detail shot hands the model a pair of
  hands and asks it to invent a woman for the next ten seconds.

#### How to write the still

Three things belong in frame one that people habitually leave for the clip:

- **Staging.** A messy bed, kicked-off shoes, a stack of clutter. A model asked
  mid-clip to mess up a tidy room will mess it up while the camera is running.
- **Text.** A video model asked to invent a headline gets close and no closer.
  An image model spells it. Put the wordmark, the timestamp and the scorebug in
  the picture and let the clip preserve them.
- **Shot size.** Frame one _is_ the opening shot, at 35 dB. Compose it as the
  first shot of the film, not as a portrait of the subject.

And one thing that belongs nowhere: **make unwanted text unreadable rather than
banned**. A wall of photo prints described as _"too small and too far out of
focus to read"_ leaked no text at all. A banned string the composition still
wants is a string the model draws badly.

**Worked examples:** [Nobody filmed this](#nobody-filmed-this) and
[Nine shots and not one cut](#nine-shots-and-not-one-cut).

### Pipeline B: animate a picture you already have

A generation already in your library, or a photograph. No second render, no
second charge — and frame one is **fixed**, rather than something you can buy
your way out of.

#### What it is for

- **Best for:** the cheapest clip there is, since the picture is already paid
  for, and for bringing back something you made weeks ago.
- **Costs you:** the lever. On a still you draw, anything the clip needs can be
  added in step one. Here the ten seconds have to be written to the picture you
  already have.
- **Watch for:** the aspect ratio. Left on Auto, the clip takes **the image's
  own shape**, which is how a 4:3 clip ends up in a 9:16 feed.

#### How to write it

Read the picture first, and write only what it can support. The ride below is
the case: the frame holds a mane, a saddle, a rein and a slice of shoulder —
**no legs and no head** — so the brief never describes a gait. It describes the
_evidence_ of one, and puts the gait itself in the camera: _"a light regular
bounce in time with the horse's stride."_

Measuring the result says it worked. The frame moves up and down far more than
side to side (2.3 px against 0.9 px per frame on average), the bounce holds a
steady **1.6–1.8 Hz** for the full ten seconds, and over 241 frames it drifts 28
px vertically on a 1280-pixel picture. A pan piles up; this does not.

> **A frame that does not contain the hard thing does not have to solve it.**

Some pictures hand you a rule for free. A selfie is one: an arm cannot change
length, so the shot size is pinned by physics rather than by a promise — which,
by the cut rule above, is an unbroken take you did not have to argue for.

**Worked example:** [A ride nobody filmed](#a-ride-nobody-filmed).

### Pipeline C: several takes, cut in post

Fifteen seconds is the ceiling and hard cuts cannot be prompted, so anything
longer, or anything genuinely edited, is **more than one generation**.

#### What it is for

- **Best for:** a montage, a sequence with real cuts, anything past fifteen
  seconds.
- **Costs you:** continuity, which no input here can carry. There is no
  reference channel and no last-frame slot.
- **Watch for:** drift between takes. This is where Grok is weakest against the
  alternatives, and it is worth saying so plainly.

#### How to chain

[Extract the final frame](/guides/extract-a-frame-from-a-video/) of one
take, hand it in as the next take's starting image, and
[join the run](/guides/join-clips-into-one-video/) at the end. Frame one
is a 35 dB lock, so the seam is genuinely tight — tighter than the same trick on
a model that treats a first frame as an anchor.

What the chain cannot do is come back to a subject after a shot they were not
in. If your piece needs one character across eight shots from eight angles, the
honest answers are a model with a reference slot —
[Seedance 2.5](/guides/seedance-2-5-prompting/) takes thirty — or an image model
holding a character sheet, with Grok animating each frame it produces.

We have no finished chain to show yet, and inventing one would be worth less
than saying so.

---

## When it breaks

### "It will not let me generate — it wants an image"

There is no text-to-video door on this model. The composer blocks the request
before it costs anything, with _"Grok Imagine Video 1.5 needs an image of the
subject — attach one before generating."_

> **Fix:** make a still first, or pick one out of your library. If you wanted a
> clip from words alone, that is
> [Seedance](/guides/seedance-2-5-prompting/) or
> [Veo](/guides/veo-3-1-prompting/) rather than this.

### "My cuts did not happen"

You cannot set the transition style from the prompt. Every ban, every
`HARD CUT.`, every _"nine shots"_ in a style line — none of it changes how a
shot change is handled.

> **Fix:** if the next shot is at a different **distance** from the one before
> it, the model cuts on its own. If it is not, no wording will buy you a cut,
> and the piece has to be rendered as separate takes and
> [joined](/guides/join-clips-into-one-video/).

### "The aspect ratio is not the one I asked for"

Left on Auto, the clip takes the **starting image's** own shape.

> **Fix:** set the ratio explicitly, and make the still in the shape you need so
> the two agree. A 4:3 picture handed to a 9:16 setting is a crop somebody has
> to decide about, and it is cheaper to decide it in step one.

### "The audio track is silent"

The brief asked for atmosphere and gave the model nothing to attach it to.
Ambience on its own comes back at −65 LUFS or below, which nobody can hear.

> **Fix:** name events, with the material they happen on — footsteps on gravel,
> a mug set down on wood, a shutter crashing, a laugh. If the shot really has
> nothing in it that makes a noise, accept the silent file and lay your own
> track under it. The audio was not charged for separately anyway.

### "The camera moved when I told it not to"

You gave it an amount. A percentage, a _"slight"_, a _"nearly invisible"_ — each
of them reads as a gesture toward movement, which is movement.

> **Fix:** state the end state instead. _"The framing at ten seconds is
> identical to the framing at zero seconds"_ produced a pixel-identical pair of
> frames over 240 frames.

### "Everything after the first beat arrived late"

There is a performance in the way. A turn, a walk, any big physical move takes
as long as a body takes, and every beat behind it slides.

> **Fix:** put the turn last, or give it a window of its own with nothing queued
> behind it. Cuts can land early; bodies cannot.

### "The last second goes slack"

Two actions in the closing beat. The model performs the first and settles into
the end.

> **Fix:** one action in the final block, and make it the one you want on screen
> at the buzzer.

### "The object in my picture turned into a different object"

The model copies the frame you gave it, then rebuilds each object from scratch
as the angle changes. A prop described in the prompt does not survive that on
its own.

> **Fix:** keep the object at a consistent angle, keep the take short, or render
> the shots that show it from a new angle as separate generations. If it is a
> graphic rather than a prop, composite it afterwards.

### "The text in my frame came back redrawn"

The words survive; the design does not. Expect the spelling to be right and the
layout to be reinterpreted, then held steady for the rest of the clip.

> **Fix:** anything whose design has to be exact is a compositing job. Anything
> whose wording has to be exact belongs in the still, large in frame.

---

## Worked examples

Three finished pieces, each with the brief that made it. Every one is a
**single generation** — nothing here is several takes cut together, which on
this model means nothing here has a cut in it unless the model chose to put one
there.

### Nobody filmed this

_Pipeline A: a drawn plate, then ten seconds of a camera that cannot move._

<figure class="guide-example">
<div class="guide-example-stage"><video src="/showcase/guides/the-night-shift.mp4" poster="/showcase/guides/the-night-shift.webp" width="608" height="1080" controls controlslist="nodownload noplaybackrate noremoteplayback" disablepictureinpicture playsinline preload="none" aria-label="Infrared doorbell footage in which a raccoon sets a parcel down on a doormat, then a dozen pairs of eyes open in the dark behind it and a dozen more raccoons walk out carrying parcels of their own"></video></div>
<figcaption class="guide-example-side">
<dl class="guide-example-spec">
<div><dt>Length</dt><dd>10s, one take</dd></div>
<div><dt>Frame</dt><dd>9:16, 720p</dd></div>
<div><dt>Supplied</dt><dd>One drawn plate</dd></div>
<div><dt>Cost</dt><dd><span class="credit">11</span> + <span class="credit">67</span></dd></div>
</dl>
<img class="guide-example-input" src="/showcase/guides/the-night-shift-first-frame.webp" width="540" height="960" loading="lazy" decoding="async" alt="The supplied plate: an infrared porch at night, a raccoon standing on the doormat holding a white parcel, camera chrome burned into the corners" />
<p class="guide-example-input-note">The plate, made first — including the burned-in camera chrome.</p>
</figcaption>
</figure>

A doorbell camera chooses nothing. It is bolted to a wall, it cannot pan, it
cannot cut, and it does not know it is filming. That makes the fixed mount the
whole genre, and one instruction matters more than all the others.

**The camera did not move once.** Frame 0 and frame 239 put the door edge, both
hinges, the panel lines, the mat edge, the step curve and the vignette on
identical pixels. That is the end-state rule at work: an earlier broadcast take
asked for a push-in of _"no more than three percent"_ and got twenty, while this
one asked for the framing at ten seconds to equal the framing at zero.

The camera chrome is in the plate rather than the prompt, and it survived
untouched — `FRONT DOOR`, the timestamp, the battery and the `REC` dot, same
style, same positions, frozen. A thin corner burn is closer to what the model
expects than a broadcast graphic, and the broadcast graphic is the one that came
back redrawn.

Two things it teaches about sound. The eye-shine gag needed no sound at all, and
the model made its own atmosphere for it — but the brief's `Audio:` line is
mostly ambience, and the track measures **−48.7 LUFS**. Audible if you go
looking, invisible in a feed. The chase below asked for a klaxon, rain,
footfalls, a shutter and a body hitting brick, and came back thirty decibels
louder.

<details class="guide-fold"><summary>The clip's brief, in full</summary><pre><code>
The frame is already correct. Animate it and change nothing else.&#10;&#10;Camera: bolted to the wall beside the door. It cannot move. The framing at&#10;ten seconds is identical to the framing at zero seconds — same crop, same&#10;angle, same lens, same distance from the mat. No pan, no tilt, no zoom, no&#10;push-in, no handheld drift, no rack focus, no cut.&#10;&#10;0-3s: the raccoon crouches and sets the white parcel down on the doormat,&#10;then nudges it with one paw until it sits square to the edge of the mat.&#10;&#10;3-4s: it rises back onto its hind legs and looks straight into the camera&#10;lens, and holds still.&#10;&#10;4-5s: behind it, in the solid black beyond the mat, pairs of small bright&#10;eyes open and catch the infrared. A dozen of them, spread across the&#10;darkness at the same low height. Nothing else of them is visible yet. The&#10;raccoon in front does not turn around.&#10;&#10;5-10s: they walk forward out of the black into the infrared light, one&#10;raccoon behind each pair of eyes, every one of them carrying a plain white&#10;parcel of its own. They reach the mat unhurried and set their parcels down,&#10;building a neat stack. The first raccoon steps aside and lets them work. At&#10;ten seconds the stack stands on the mat and several of them are looking&#10;directly into the lens.&#10;&#10;Light: constant infrared, blown out near the door, solid black past the mat.&#10;No exposure shift, no flicker, no colour anywhere, no porch light switching&#10;on.&#10;&#10;Overlay: FRONT DOOR, the timestamp and the REC dot are burned into the&#10;frame. They do not move, drift, warp or re-spell.&#10;&#10;Audio: night ambience, faint crickets and a distant road. Claws on concrete&#10;and the scuff of cardboard as the first parcel is set down. A beat of near&#10;silence as the eyes open. Then many sets of claws on concrete at once, and&#10;cardboard settling on cardboard as the stack builds. No music, no voices, no&#10;dialogue.&#10;&#10;Avoid: camera movement of any kind, cuts, dissolves, speed ramps, a raccoon&#10;that morphs or changes size, a human appearing, the door opening or closing&#10;further, colour appearing in the frame, on-screen text other than the&#10;overlay already in the frame, captions, watermarks, a doorbell chime.</code></pre><p class="guide-fold-note">And the plate it animates, made first on GPT Image 2. Every word the clip has to keep is written here, not in the clip's brief.</p><pre><code>
A frozen frame from a home doorbell camera at night, infrared night vision.&#10;&#10;Look: monochrome with a cold blue-grey cast and no colour anywhere. A harsh&#10;infrared floodlight from the camera itself blows out the doorframe and the&#10;near concrete to almost white, and falls off to solid black about two feet&#10;past the doormat. Wide fisheye lens, strong barrel distortion bending the&#10;edge of the step into a curve, dark vignetted corners. Soft focus, visible&#10;sensor noise, heavy compression, the flat cheap look of real security&#10;footage.&#10;&#10;Framing: vertical, the camera fixed to the wall beside the door at chest&#10;height and angled slightly down onto a concrete stoop. On the left a white&#10;panelled front door stands half open, its small window catching the&#10;infrared. Centre, a dark rectangular doormat on the concrete. Beyond the&#10;mat, total blackness.&#10;&#10;Subject: a raccoon standing upright on its hind legs on the doormat, facing&#10;the camera dead-on, holding a small plain white cardboard parcel against its&#10;chest with both front paws. Thick fur, ringed tail, black mask across the&#10;eyes, eyes bright and reflective in the infrared. Lightly stylised, a&#10;character from a 3D animated film rather than a photograph of a wild animal.&#10;Calm and deliberate, as if making a delivery.&#10;&#10;Overlay burned into the frame, thin white sans-serif, crisp and correctly&#10;spelled. Top left: FRONT DOOR. Top right: 08/05/2026 02:47:13. Bottom left:&#10;a small battery icon, then a small filled dot, then REC.&#10;&#10;No other text, no people, no other animals, nothing else on the porch.</code></pre></details>

### A ride nobody filmed

_Pipeline B: a picture that already existed._

<figure class="guide-example">
<div class="guide-example-stage"><video src="/showcase/guides/the-ride.mp4" poster="/showcase/guides/the-ride.webp" width="608" height="1080" controls controlslist="nodownload noplaybackrate noremoteplayback" disablepictureinpicture playsinline preload="none" aria-label="A woman rides across Mongolian grassland filming herself at arm's length, laughs into her own lens, then throws her head back into the wind"></video></div>
<figcaption class="guide-example-side">
<dl class="guide-example-spec">
<div><dt>Length</dt><dd>10s, one take</dd></div>
<div><dt>Frame</dt><dd>9:16, 720p</dd></div>
<div><dt>Supplied</dt><dd>One existing picture</dd></div>
<div><dt>Cost</dt><dd><span class="credit">0</span> + <span class="credit">67</span></dd></div>
</dl>
<img class="guide-example-input" src="/showcase/guides/the-ride-first-frame.webp" width="540" height="960" loading="lazy" decoding="async" alt="The starting picture: the same rider on horseback in a brown embroidered deel, arm out of frame at the bottom left, grassland and two white gers behind her" />
<p class="guide-example-input-note">Already in the library, already paid for. Nothing was uploaded.</p>
</figcaption>
</figure>

The cheapest clip on this page, because frame one was made weeks earlier for
something else. **Sixty-seven credits, first submission, no re-rolls.**

The picture decides almost everything. Her arm leaves the frame at the bottom
left, so the camera is a hand on a moving horse — and of the horse, the frame
holds a mane, a saddle, a rein and a slice of shoulder, **no legs and no head**.
So the brief never asks for a gait. It asks for the mane to lift, the hair to
stream, the tassels to swing and the ridge to slide past, and puts the trot
itself in the camera as _"a light regular bounce in time with the horse's
stride."_ The bounce came back steady at 1.6–1.8 Hz for the full ten seconds.

**The audio is why this one is worth watching with the sound on.** Three earlier
takes asked for room tone under a silent performer and came back silent. This
one asked for a laugh, and the laugh arrives at **−26.9 LUFS**, forty decibels
above them. It also stops dead at 6.40s, the frame her head turns away from the
lens, while she is still visibly laughing for another two seconds.

And the thing nobody wrote: _"looks off past the camera toward the hills for a
moment"_ came back as a head thrown fully back into the wind, hair streaming,
held for two seconds. That is the standard image of this format, and both halves
of the instruction — _toward the hills_, _for a moment_ — lost to it.

<details class="guide-fold"><summary>The brief, in full</summary><p class="guide-fold-note">Three beats in ten seconds, deliberately wide windows, and not one sentence describing the light, the clothes or the land she is already standing on.</p><pre><code>
The woman in the first frame is riding a horse at a steady trot across open&#10;Mongolian grassland, filming herself on her phone held out at arm's length.&#10;One continuous unbroken take, ten seconds, no cuts and no transitions of any&#10;kind.&#10;&#10;She is the same woman the whole way through — same face, same long dark&#10;wind-blown hair, same dark brown deel with gold-brown embroidered trim, same&#10;wide silver belt with the round silver medallion and its amber stone, same&#10;pair of silver tassel pendants hanging at her hip. Her right hand keeps hold&#10;of the rein at the red saddle the whole time. Nothing about her drifts,&#10;morphs or re-rolls.&#10;&#10;The camera is her own hand at the end of her outstretched arm, so it rides&#10;with her: a light regular bounce in time with the horse's stride, a small&#10;amount of hand sway, nothing more. It never pans, tilts, zooms, pushes in or&#10;pulls back, and it never leaves her hand. Her arm's length never changes, so&#10;she is the same size in the frame in the last second as in the first,&#10;filling the frame from the waist up and slightly off centre.&#10;&#10;Show that the horse is moving through the horizon rather than through its&#10;legs: the black mane in the lower right lifts and falls in time with the&#10;bounce, her hair streams back off her shoulders, the tassels at her belt&#10;swing, and the grassland, the fence and the two white gers behind her drift&#10;steadily past. Keep the horse's legs and head out of the frame exactly as&#10;they are now.&#10;&#10;0.0s — She holds the frame, riding, smiling into the lens as her hair blows&#10;across her face and she lets it.&#10;&#10;3.0s — She laughs — a real open laugh with her eyes creasing and her&#10;shoulders shaking, chin lifting a little, delighted. She keeps her eyes on&#10;the lens and keeps hold of the rein.&#10;&#10;6.5s — Still laughing, she looks off past the camera toward the hills for a&#10;moment, then back to the lens, and the laugh eases down into a wide open&#10;grin as she rides on.&#10;&#10;Land: rolling green-brown steppe under a low overcast sky, layered grey&#10;cloud, dark hills along the horizon, two white gers and a wooden fence on&#10;the ridge behind her, flat soft daylight with no sun on her face. The&#10;weather does not change.&#10;&#10;Sound: her laughing, wind across an open plain, the muffled rhythm of hooves&#10;on grass, and the creak of a saddle. No music, no voiceover, no other&#10;people, no speech — she laughs but she never says a word and never mouths&#10;one.&#10;&#10;Style: photorealistic phone selfie footage shot at arm's length from&#10;horseback, real handheld bounce, real blinking, breath in her shoulders,&#10;hair with weight to it, fine sensor grain, 24fps.&#10;&#10;Avoid: cut, hard cut, jump cut, dissolve, cross-fade, camera move, zoom,&#10;pan, push in, pull back, a change of shot size, the camera leaving her hand,&#10;a drone shot, a wide landscape shot, the horse's legs, the horse's head, a&#10;horse rearing or bolting, a second rider, a second horse, another person, a&#10;hat, dropping the rein, a moving mouth forming words, text overlay,&#10;subtitles, captions, watermark, logo, timestamp, readable writing anywhere&#10;in frame.</code></pre></details>

### Nine shots and not one cut

_Pipeline A at full length, and the clearest example on this page of what the
model will not do._

<figure class="guide-example">
<div class="guide-example-stage"><video src="/showcase/guides/the-chase.mp4" poster="/showcase/guides/the-chase.webp" width="608" height="1080" controls controlslist="nodownload noplaybackrate noremoteplayback" disablepictureinpicture playsinline preload="none" aria-label="A hacker runs a stolen machine out of a building through a rain-soaked neon city: a server aisle, a stairwell drop, an alley, a loading dock, a rooftop leap and a ledge catch"></video></div>
<figcaption class="guide-example-side">
<dl class="guide-example-spec">
<div><dt>Length</dt><dd>15s, one take</dd></div>
<div><dt>Frame</dt><dd>9:16, 720p</dd></div>
<div><dt>Supplied</dt><dd>One drawn first frame</dd></div>
<div><dt>Cost</dt><dd><span class="credit">4</span> + <span class="credit">100</span></dd></div>
</dl>
<img class="guide-example-input" src="/showcase/guides/the-chase-first-frame.webp" width="540" height="960" loading="lazy" decoding="async" alt="The supplied first frame: the character sprinting at the camera down a dark server aisle under a red alarm strobe" />
<p class="guide-example-input-note">Frame one has to hold a face, so the film opens on a sprint at the lens.</p>
</figcaption>
</figure>

Fifteen seconds, nine shots, a hundred credits — and **zero cuts at every
detection threshold from 0.3 down to 0.1**. Not fewer than asked: none. All nine
shot changes are cross-fades, each about a third of a second long.

The beats are all there and all in order: the server aisle, the deck close-up,
the stairwell, the puddle, the searchlit alley, the shutter, the rooftop leap,
the ledge catch, the kneel. The brief asked for hard cuts **eight times** —
three `HARD CUT.`s inside the beats, once in the style line, once in `Avoid:`.
The same brief on Seedance cut eight times. This is one clip showing both the
negation rule and the cut rule.

Two more findings sit in it. **The prop drifted while the pixels held**: frame
one has a flat matte-black deck with a hot-pink X on the lid, and by the second
second she is carrying a red-cased smartphone. And **the ending slows down** —
the closing beat asked her to pocket the deck _and_ push up out of the crouch to
run again, and the last 1.4 seconds run 14.5 → 4.4 on frame-to-frame
difference. It did the first half.

The one thing that held completely is what the run was built to test. Fifteen
seconds, nine shots, **no reference channel at all**, and her face, the
ponytail, the white X on the hoodie, the hip chain and the pink-soled high-tops
are all still recognisable at fourteen seconds. **A first frame did a reference
channel's job for a quarter of a minute.**

<details class="guide-fold"><summary>The brief, in full</summary><p class="guide-fold-note">Reproduced exactly as it was submitted, including the eight cut instructions that did nothing. Six timestamped windows carry nine shots, and all nine shot changes came out as cross-fades.</p><pre><code>
Style: photorealistic cinematic night action, handheld and urgent, in a&#10;rain-soaked neon city. Hard fast cutting — NINE shots in fifteen seconds,&#10;every cut a hard cut, no dissolves or fades. Deep blacks, wet reflective&#10;concrete, hot magenta and cyan sign light, cold blue shadow, heavy rain,&#10;steam, sparks. Real motion blur on the fast whips, restrained grain.&#10;Vertical 9:16.&#10;&#10;Subject: JANA — the young woman in the opening frame, identical in every&#10;shot: blonde hair with hot-pink streaks in a high messy ponytail, pale skin,&#10;grey-violet eyes, a small inked black X with three dots under her left eye,&#10;a black studded choker, an oversized black hoodie with a white brushed X on&#10;the chest, black cargo trousers with grey straps and a hanging chain, chunky&#10;black-white-and-pink high-top sneakers. Her face, hair, marking and outfit&#10;never drift, morph or re-roll. She is moving hard in every shot and never&#10;stands still. She never speaks.&#10;&#10;Prop: THE DECK — a flat matte-black handheld computer with a hot-pink X on&#10;the lid and a glowing pink edge strip. It is in her hand or on a strap&#10;across her body in every shot.&#10;&#10;Action Sequence — nine shots, hard cuts:&#10;  0.0s - 2.5s    Dark server aisle, red alarm strobes. Camera racing&#10;    backwards ahead of her: JANA sprints straight at the lens in the black&#10;    hoodie, deck in one hand, ponytail thrown up by the movement. HARD CUT.&#10;    MACRO low on her hands: her thumb drags across the deck's glowing pink&#10;    screen as she runs, a red strobe snapping across her knuckles.&#10;  2.5s - 5.0s    Concrete stairwell, camera at the bottom of the well&#10;    looking straight up. JANA vaults the handrail two floors up, drops a full&#10;    flight through frame with the deck clamped to her chest, and slams onto&#10;    the landing in a crouch, sneakers skidding.&#10;  5.0s - 7.5s    Alley, hammering rain. GROUND-LEVEL MACRO on her&#10;    black-white-and-pink sneakers smashing through a magenta-lit puddle. HARD&#10;    CUT. Camera racing backwards ahead of her: she sprints into the lens, hood&#10;    up, breath fogging, searchlights swinging over wet brick behind her.&#10;  7.5s - 10.0s   Loading dock, camera flat on the wet ground. A steel roller&#10;    shutter drops fast; JANA slides under it on her hip with the deck hugged&#10;    to her chest, and it crashes shut behind her sneakers.&#10;  10.0s - 12.5s  Rooftop, city glow below. Tracking side-on: she runs the&#10;    wet parapet in the black hoodie and launches over the gap between two&#10;    buildings, arms wide, ponytail streaming. HARD CUT. TIGHT on her hand&#10;    slapping the far ledge and gripping, her body slamming into wet brick.&#10;  12.5s - 15.0s  Lower roof, rain through neon, camera low and close. She&#10;    hauls herself over the ledge onto one knee, flips the deck open and&#10;    hot-pink screen light hits her face and the inked X under her eye. She&#10;    looks up past the lens, grins hard, shoves the deck into her hoodie pocket&#10;    and pushes up out of the crouch to run again.&#10;&#10;Audio: an alarm klaxon and server fans; her hard breathing throughout;&#10;sneakers on concrete, steel and wet asphalt; rain hammering metal; a steel&#10;shutter crashing down; wind at the roof edge; a body hitting brick; the deck&#10;snapping open. No music, no score, no narration, no dialogue, nobody&#10;speaking.&#10;&#10;Avoid: no on-screen text, subtitles, captions, watermark, logo, HUD or&#10;interface overlay; no character sheet, model sheet, turnaround, multiple&#10;panels, side-by-side figures, split screen, white studio background,&#10;watercolour splash or colour swatches; no standing still, posing or static&#10;shot; no slow motion, speed ramp or freeze frame; no dissolve, fade or wipe;&#10;no guns, gunfire, blood or injury; no second woman resembling her; no change&#10;of hair colour, hairstyle or clothing; no visible camera or crew; no&#10;morphing or drifting of her face, eye marking or clothes.</code></pre></details>

---

## About justpictur.ing

justpictur.ing is an AI picture studio: prompt-driven images, clips, voice-overs and music, and storyboard-backed films with consistent characters across every scene.

Agents connect over MCP and can generate directly:

```
claude mcp add justpicturing https://api.justpictur.ing/mcp --transport http
```

The app is at https://app.justpictur.ing and requires an account.

Every guide on this site has a Markdown twin at its own URL plus `.md`. Every guide: https://justpictur.ing/guides/
