---
title: "Veo 3.1: the complete prompting guide"
description: "Eight seconds of video that writes its own soundtrack. Three audio channels, eight rules, four pipelines and the failures worth knowing about first."
canonical_url: "https://justpictur.ing/guides/veo-3-1-prompting/"
published: 2026-08-14
updated: 2026-08-16
locale: "en"
models: ["veo-3.1", "veo-3.1-fast", "veo-3.1-lite", "seedream-5-pro", "gpt-image-2"]
---

Veo 3.1 makes **eight seconds of video with its own soundtrack**: the line
spoken in the voice you described, the footsteps under it, the room it was
recorded in, and a score if you let it. All of it comes out of the same prompt
as the picture, written in three channels the model reads separately.

**Veo 3.1 is three models**, and they take the same brief.

- **Veo 3.1** takes reference images.
- **Veo 3.1 Fast** is the same model at lower fidelity. Perfect for rehearsal.
- **Veo 3.1 Lite** cannot turn its audio off. Perfect for high volume.

**NOTE:** Everything in this guide is true of all three unless it names one.

## Overview

### The short version

Four decisions do most of the work. If you read nothing else, read these. Each
one is written out in full further down.

<ol class="guide-keys">
<li class="guide-key">
<p class="guide-key-title">Write the sound in its own three channels</p>
<p class="guide-key-body">Veo reads quoted speech, <code>SFX:</code> and <code>Ambient noise:</code> as three separate things. <strong>A sound buried in a sentence about the picture gets you the picture</strong>: describing a knife "with a satisfying crunch" gets you a knife. A line that begins <code>SFX:</code> gets you the crunch.</p>
<div class="guide-key-lines">
<div class="guide-key-quote guide-key-quote--no"><p class="guide-key-quote-label">Instead of</p><p class="guide-key-quote-text">The blade presses down with a satisfying crunch</p></div>
<div class="guide-key-quote guide-key-quote--yes"><p class="guide-key-quote-label">Write</p><p class="guide-key-quote-text">SFX: a dense crystalline crunch, releasing as the blade breaks through</p></div>
</div>
</li>
<li class="guide-key">
<p class="guide-key-title">Give the last block an action, never a decay</p>
<p class="guide-key-body">Veo front-loads a <span class="term"><button class="term-word" type="button" popovertarget="def-beat-sheet">beat sheet</button><span class="term-def" popover id="def-beat-sheet">A list of what happens when: one line per moment, each with the seconds it occupies. It is the shape of the film written down before anything renders.</span></span> and will not stretch anything to fill a tail. A final block naming a pose, a silence or a sound dying away stops the picture and the sound together — <strong>44% of one eight-second take came back below −41 dB</strong> because its last line asked for a ring to fade.</p>
<div class="guide-key-lines">
<p class="guide-key-practice">Name a real event in the final block: a slice rocking to rest, a knife set down, a last line delivered on the way out. The take whose blocks all named actions had no gap anywhere in it.</p>
</div>
</li>
<li class="guide-key">
<p class="guide-key-title">Name the rig, not the shot</p>
<p class="guide-key-body">"POV shot" and "selfie angle" ask for a framing and get half of one. <strong>A camera stated as an object in the world — a thing, in a named hand — gets followed</strong>, and the framing falls out of it. Every working version of the vlog format says some variant of this sentence.</p>
<div class="guide-key-lines">
<div class="guide-key-quote guide-key-quote--no"><p class="guide-key-quote-label">Instead of</p><p class="guide-key-quote-text">POV selfie shot, wide angle</p></div>
<div class="guide-key-quote guide-key-quote--yes"><p class="guide-key-quote-label">Write</p><p class="guide-key-quote-text">He holds a selfie stick in his right hand, at arm's length — that is where the camera is</p></div>
</div>
</li>
<li class="guide-key">
<p class="guide-key-title">Reference images lock the ratio and the length</p>
<p class="guide-key-body">Reference-to-video runs at <strong>16:9 and eight seconds, and at nothing else</strong>, and any <span class="term"><button class="term-word" type="button" popovertarget="def-last-frame">last frame</button><span class="term-def" popover id="def-last-frame">A picture you supply that the clip has to arrive at. It becomes the final frame, and the model invents the motion that gets there.</span></span> you supply alongside is dropped. Veo 3.1 Fast and Veo 3.1 Lite have no reference slot at all.</p>
<div class="guide-key-lines">
<p class="guide-key-practice">For anything vertical, describe the subject in prose or start from a first frame in the shape you need. There is no vertical reference door, on any model in the family.</p>
</div>
</li>
</ol>

### The clock is three numbers

Most video models take a duration the way an editor does: any length inside a
range, and often `auto` if you would rather not decide. Seedance 2.5 runs
anywhere from 4 to 30 seconds, or picks for itself.

Veo does neither. **Four, six or eight seconds** — three numbers, nothing between
them, nothing past eight, and no auto.

So length is not a setting here, it is the structure. You write to a clock rather
than to a character count: **four blocks of two seconds** at the ceiling, and
anything longer is two <span class="term"><button class="term-word" type="button" popovertarget="def-take">takes</button><span class="term-def" popover id="def-take">One run of the model: a single generated clip, start to finish. Running the same brief again gives you another take, not a revision of this one.</span></span>
[joined end to end](/guides/join-clips-into-one-video/).

### What it accepts

| Input            | Limit             | Notes                                                                                     |
| ---------------- | ----------------- | ----------------------------------------------------------------------------------------- |
| Prompt           | 4,000 chars       | Roomy. What binds here is the clock, not the character count.                             |
| Duration         | 4, 6 or 8 seconds | Three steps, not a slider, and eight is the ceiling.                                      |
| Resolution       | 720p · 1080p      | Both, on all three tiers. On Lite, 1080p only runs at eight seconds.                      |
| Aspect ratio     | 16:9 or 9:16      | Two, and there is no third. Not 1:1, not 4:5, not 21:9.                                   |
| First frame      | 1 image           | The clip opens on it exactly.                                                             |
| Last frame       | 1 image           | The clip arrives at it. Needs a first frame, and is **ignored if you supply references**. |
| Reference images | up to 3           | **Veo 3.1 only. 16:9 and 8 seconds only.**                                                |
| Audio            | a switch          | On Veo 3.1 and Fast. Veo 3.1 Lite has no switch and always makes sound.                   |

### What it costs

**Veo 3.1 and Veo 3.1 Fast price on the audio switch and on nothing else**, so a
1080p take and a 720p one come to the same number and there is no reason to shoot
the smaller.

| Length | Veo 3.1 silent                    | Veo 3.1 + audio                   | Fast silent                      | Fast + audio                    |
| ------ | --------------------------------- | --------------------------------- | -------------------------------- | ------------------------------- |
| 4s     | <span class="credit">67</span>    | <span class="credit">133.5</span> | <span class="credit">33.5</span> | <span class="credit">50</span>  |
| 6s     | <span class="credit">100</span>   | <span class="credit">200</span>   | <span class="credit">50</span>   | <span class="credit">75</span>  |
| 8s     | <span class="credit">133.5</span> | <span class="credit">267</span>   | <span class="credit">67</span>   | <span class="credit">100</span> |

**Veo 3.1 Lite prices on resolution instead**, and it has no switch: every Lite
take generates audio, so there is no silent column and no way to make one. At
1080p it only runs at eight seconds.

| Length | 720p                             | 1080p                            |
| ------ | -------------------------------- | -------------------------------- |
| 4s     | <span class="credit">17</span>   | —                                |
| 6s     | <span class="credit">25</span>   | —                                |
| 8s     | <span class="credit">33.5</span> | <span class="credit">53.5</span> |

<blockquote class="guide-quote--caution">
<p>On the full model, <strong>sound doubles the rate</strong>. Turn it off for anything with nothing to say — a landscape, a texture, a plate you intend to score yourself — and leave it on the moment a person opens their mouth.</p>
</blockquote>

---

## The rules

Eight things that are true whatever you are making. Each one came out of a take
that went wrong, and they are ordered by how often they matter.

### Write the sound in its own three channels

> Veo separates them. A prompt that does not is asking one channel to carry three.

Google's own notation is three marks, and the model reads each of them as its
own track:

```
"…"              dialogue, quoted verbatim
SFX:             events, on their own line, attached to the beat they belong to
Ambient noise:   the bed that runs under all of it
```

Burying _"with a satisfying crunch"_ inside a sentence about a knife gets you a
knife. A line that begins `SFX:` gets you the crunch. Every audio beat in the
[glass strawberry brief](#two-cuts-eight-seconds) is on its own `SFX:` line
attached to its own timestamp, and the levels came back landing where they were
asked to.

Prose works. It just works less, for the same number of characters.

### Give the last block an action, never a decay

> The model plays to the buzzer only if the buzzer has something happening on it.

Veo front-loads a beat sheet, and it will not stretch action to fill a tail it
was told nothing happens in. This is measurable.

- A final block reading _"the last of the glass ringing decaying away into
  silence"_ produced **3.5 seconds of dead air** — 44% of the runtime, every peak
  below −41 dB, and the picture frozen with it.
- A final block naming a wipe, a face coming back and a line delivered came back
  with **no gap anywhere**, mean level between −11.2 and −19.5 dB across all
  sixteen half-second windows.

Same model, same length, same week. The difference is that one last block named
an absence and the other named an event.

### A timestamp is a budget, and Veo keeps it

> Four blocks of two seconds is the working unit at eight.

Writing `[00:00-00:02]` is not a hint. In a four-block sheet the four spoken
lines landed **four for four**, each inside its own window, in order, verbatim.

The strongest evidence that it reads the sheet rather than the words is the
silence: a block that named an action and no dialogue produced **2.1 seconds of
deliberate quiet**, and the model came back in on time afterwards.

Two seconds is about as fine as it is worth cutting. Below that you are dividing
eight seconds into pieces too small to hold an event, and what you get is a
rushed one.

### Name the rig, not the shot

> A shot type asks for a framing. An object asks for a thing, and the framing follows.

_"POV shot"_ and _"selfie angle"_ both underperform. `"He is holding a selfie
stick in his right hand, at arm's length, a little above his own eyeline and
angled down at him — that is where the camera is"` does not.

It works twice over. The reference still came back with his arm already running
out of frame toward the lens, so the rig existed before the clip was ever
submitted; and the gag in the final block lands because the model had already
been told which hand was busy holding the stick.

The same rule is why you can direct the camera to _change_. You cannot put down a
shot type, and you can put down a phone.

### Rehearse on Fast, finish on Veo 3.1

> Same clock, same ratios, same three channels. One of them is for finding out.

Veo 3.1 Fast is the same family and takes the same brief. It has the same 4/6/8
steps, the same two aspect ratios, the same audio switch and the same notation,
so a sheet that works on one works on the other. What it gives up is fidelity —
and reference images, which it does not have at all.

So run the brief on Fast first and read it honestly. It answers everything a
sheet can get wrong: whether the blocks land where you asked, whether the lines
fit the clock, whether the sound arrives on the right second, whether the
performance is the one you pictured. Then re-run the same words on Veo 3.1 for
the version you keep.

The one thing it cannot rehearse is a reference set, because it has no slot to
put one in. If your piece needs references, rehearse at four seconds on the full
model instead.

### A reference carries pose, not just identity

> A property the input sets cannot be argued out of the model downstream.

The reference for the vlog was a standing portrait. The brief asks for the walk
in the shot line, in the subject paragraph and inside every one of its four
blocks — _already moving at second 0_, _walking for the whole shot and never
stops_, _already mid-stride_, _keeps walking_, _does not stop walking_, _still
walking_. **He stands.** Measured off frames at 0.2s, 2.4s, 3.6s, 7.4s and 7.9s:
the same ridge crest, the same background geometry, no forward travel and no
stride cadence.

Six sentences of prose lost to one property of a picture. The fix is not a
seventh — it is to shoot the reference in the state you want: mid-stride, weight
forward, one shoulder dropped, the ridge sliding past behind him.

**Describe the motion in the picture, not in the paragraph.**

### References lock the ratio, the length and your last frame

> One door or the other, and the reference door is the narrow one.

Reference-to-video is a mode, not a setting you add to the one you already have.
Supply between one and three reference images and the request is bound to a
single shape:

- **16:9** — references run at no other ratio;
- **8 seconds** — that is also the ceiling, so what you give up is 4 and 6, not
  length;
- any **last frame is dropped**.

They are also on the full model alone — Veo 3.1 Fast and Veo 3.1 Lite have no
reference slot, so a multi-shot sequence planned around either of them will not
hold together.

That makes the choice concrete rather than aesthetic. A vertical piece with a
recurring character has no reference route: describe them in prose, or settle
them in a first frame and animate that instead.

### Veo will animate a photographed face

> The refusal you hit elsewhere is not here, and that is often the whole reason to be here.

Hand Seedance a photorealistic close-up of a person and it declines: `E005 — the
input or output was flagged as sensitive`, through every door, reference image
and first frame alike. Veo 3.1 Fast took the identical frame without comment, and
it is the stronger speech model besides.

That is the axis on which the two models actually differ for
<span class="term"><button class="term-word" type="button" popovertarget="def-ugc">UGC</button><span class="term-def" popover id="def-ugc">User-generated content: an ad shot to look like a real customer filmed it on their phone, rather than like an ad.</span></span>
work. If your piece needs a believable person talking to camera from a picture
you already have, this is the model that will do it.

It is not a licence. A face you have no right to use is still a face you have no
right to use, and the person in every example here was generated.

---

## The pipelines

Veo has four doors, and which one you use decides more about the result than
anything you write. Each excludes something.

### The blocks every brief uses

> Write it like a shooting script with a sound sheet attached. That is what it reads.

Whichever door you came through, the prompt is built from the same parts, in the
order the model reads most reliably:

```
[Shot]           the take, its length, the camera, and that it is unbroken
[Subject]        who or what must stay identical from first frame to last
[Context]        the room, the weather, the light — skipped if a frame carries it
[Action]         the beat sheet, one block per two seconds
  SFX:             the events inside that block, on their own line
[Ambient noise]  the bed that runs under all of it
[Style]          the stock, the lens, the grade
[Avoid]          what must never appear
```

There is **no negative-prompt field** here, so the `Avoid:` line at the foot of
the brief is where exclusions live. It is read: every brief in the worked
examples below ends with one, and the things they ban stayed out.

At 4,000 characters for eight seconds of film, room is not what you are short
of. What you are fighting for is attention.

> **The rule for cutting: delete anything that cannot happen in eight seconds.**
> A wardrobe change, a second location, a character arc — none of them fit, and a
> sentence describing one is a sentence spent buying nothing.

### How to write the audio block

Speech is quoted verbatim inside the beat it belongs to. Events get an `SFX:`
line of their own. The bed gets one `Ambient noise:` line for the whole clip.

```
[00:02-00:04] He keeps walking, his breath fogging past the lens, and says,
"Forty minutes back." The wind rises behind him.
SFX: fur and wind buffeting the microphone, a low gust building.

Ambient noise: a constant high-altitude wind under everything, thin and cold.
```

Three things are worth knowing before you write one.

**Quote the line; do not describe it.** `she says, "That's the part nobody tells
you."` gets you that line. `she says something encouraging` gets you a mouth
moving around syllables that are not words, because you asked for a description
and the model rendered the description.

**Count the words against the clock.** Eight seconds is roughly twenty words at a
natural pace, across every block combined. Write twenty-five and the take ends
mid-word.

**Say "no music" if you mean it.** Left alone Veo scores a scene more often than
not, usually a low pad, and it arrives already mixed under the dialogue where
nothing can separate it afterwards. If you do want a score, name the instrument
and the feel — _solo piano, sparse, one note at a time_ — because a genre name
gets you a trailer.

### Pipeline A: start from text

Nothing but a written brief.

#### What it is for

- **Best for:** <span class="term"><button class="term-word" type="button" popovertarget="def-b-roll">b-roll</button><span class="term-def" popover id="def-b-roll">Footage with nobody speaking in it, like scenery, hands, a street or a detail, cut in around the main shots to cover a voice-over or to let a scene breathe.</span></span>, <span class="term"><button class="term-word" type="button" popovertarget="def-establishing">establishing shots</button><span class="term-def" popover id="def-establishing">The wide shot that opens a scene and tells you where you are before anyone in it does anything.</span></span>, atmosphere, a single spoken line in a place you do not need to match.
- **Costs you:** control. You get a good version of your description, not your
  specific person or product.
- **Watch for:** it is the only door where nothing is holding a face or an object
  in place across the eight seconds.

#### How to write it

Every block has to work, because nothing is supplied. Spend the characters on the
**subject** and the **beats**, and write the look as camera facts rather than
mood words: a named lens, a stated light direction, a grain. _"Cinematic"_ buys
nothing.

Describe anyone who has to stay the same **once**, in detail, and then point back
to that description rather than repeating it. Repeating invites a redraw.

### Pipeline B: start from text with a first frame

You have one picture, and the clip begins on it exactly.

#### What it is for

- **Best for:** bringing a settled picture to life — a product shot, a poster, a
  character, a screen.
- **Costs you:** reference images. It is one road or the other.
- **Watch for:** it is the door that keeps 9:16, which references do not.

#### How to write it

> This is the pipeline where **less prompt is better.**

The frame already carries the framing, the clothes, the room, the light and the
colour. Describing them again wastes characters and risks contradicting the
picture — and if your words and the picture disagree, you have made something
uncertain that was already settled.

Say _"the woman of the first frame, unchanged"_ once, then spend everything else
on what **changes**: the motion, the performance, the speech, the camera.

#### The frame is a gate

Make the still, then look at it before anything moves. At eight seconds the frame
is not there to hold the shot together — drift is a small risk over a clip this
short. It is there so a step-one mistake is caught while it is still one picture
and not eight seconds of motion built on top of it.

The glass strawberry is the case. _Is this glass, or is it fruit?_ is settled in
the still or it is not settled at all: a video model that starts from a real
strawberry spends all eight seconds cutting a real strawberry, and no amount of
audio description rescues it.

**Worked examples:** [Two cuts, eight seconds](#two-cuts-eight-seconds) and
[The face Seedance would not take](#the-face-seedance-would-not-take).

### Pipeline C: start from text with a first and a last frame

You have both ends, and the model invents the middle.

#### What it is for

- **Best for:** a transition you have drawn both sides of, a move that has to end
  somewhere precise, a shot that has to hand off to the next one.
- **Costs you:** reference images, which silently switch it off.
- **Watch for:** the two frames have to be a plausible eight seconds apart. It is
  an <span class="term"><button class="term-word" type="button" popovertarget="def-interpolation">interpolation</button><span class="term-def" popover id="def-interpolation">Filling in the motion between two pictures you already have, rather than inventing where a shot ends. You supply both ends; the model makes the journey.</span></span>, not a teleport.

#### How to write it

Both frames are senior to the prose. A beat that contradicts the last frame is
two instructions pulling against each other, and the picture usually wins — so
read the last frame before you write the last beat, and rewrite the beat to
describe what the picture actually shows.

This is also the one door that is the same width on all three tiers. Fast and
Lite take a first frame and a last frame exactly as the full model does, so an
interpolation is the one thing Veo 3.1 Lite does as well as anything else in the
family.

The practical argument for the second door is that it pins the ending. A beat
sheet is otherwise the only thing holding a clip's final second, and a schedule
slips the moment something in the shot physically has to travel. Two ends means
the schedule can slip and the destination cannot.

This is also how you chain past eight seconds:
[extract the final frame](/guides/extract-a-frame-from-a-video/) of one
take, hand it in as the next take's first frame, and
[join the run](/guides/join-clips-into-one-video/) at the end.

### Pipeline D: start from text with references

You have a subject — a character, a creature, a product — and it must survive the
clip.

#### What it is for

- **Best for:** a character who comes back across several clips, a face that has
  to match something already rendered.
- **Costs you:** vertical, the choice of 4 or 6 seconds, and your last frame.
  Also Fast and Lite, which have no reference slot.
- **Watch for:** the reference sets the pose as well as the identity.

#### One to three, and each one has a job

Three slots is a small budget compared with what a long-take model gives you, so
each one has to earn its place. Say what it is for, and rule out what it is not:

```
[Image1] is <subject>. Use it for <face, build, colour> only.
         Do not use its background, lighting, framing or pose.
```

Choose clear, well-lit pictures that show the subject from the angle you want it
seen from — and remember that "from the angle you want it seen from" includes the
_attitude_ you want it in, because the pose travels with it.

**Worked example:** [A commute nobody filmed](#a-commute-nobody-filmed): one
reference, eight seconds, and a beat sheet obeyed four for four.

---

## When it breaks

### "The last two seconds are dead"

The final block named a decay, a pose or a silence, and Veo does not stretch
action to fill a tail it was told nothing happens in.

> **Fix:** give the last block a real event — a slice settling, a knife set down,
> a line delivered on the way out of frame. Check the level at second seven, not
> at second one.

### "The take ends mid-word"

More words than the clock holds. Eight seconds is about twenty words at a natural
pace, across every block combined.

> **Fix:** count the words, not the lines. Cut a beat rather than speeding anyone
> up, or split the scene across two takes and join them.

### "There is music under it and I never asked for one"

Left alone, Veo scores a scene more often than not, and it arrives already mixed
under the dialogue.

> **Fix:** put `No music, no score, no soundtrack.` in the `Avoid:` line of every
> brief you intend to lay a track under later. There is no way to separate it
> afterwards, so this is a decision you only get to make once per take.

### "I cannot get a vertical clip with my character in it"

Reference images run at 16:9 and at no other ratio, on every model in the family
that has them — which is only the full one. Vertical and references cannot be
combined, so the ratio is decided the moment you reach for a reference.

> **Fix:** use a first frame in the shape you need, or describe the subject in
> prose. A first frame keeps 9:16; a reference set never does.

### "My last frame did nothing"

Reference images were supplied in the same request, and a last frame is discarded
whenever they are.

> **Fix:** pick one door. A first-and-last-frame take carries no references, and a
> reference take has no ending you can pin.

### "The character will not do the thing I keep asking for"

A reference image carries pose as well as identity, and a property the picture
sets cannot be argued out of the model in prose.

> **Fix:** re-shoot the reference in the state you want. One still beats six
> sentences, and no number of sentences wins.

### "It cost twice what I expected"

The audio switch. It doubles the per-second rate on Veo 3.1 and adds half again
on Fast.

> **Fix:** turn audio off for anything with nothing to say — a landscape, a
> texture, a plate you will score yourself. Note that Veo 3.1 Lite has no switch
> to turn off, so a silent clip has to come from Veo 3.1 or from Fast.

### "There is nowhere to put a negative prompt"

There is no field for it here.

> **Fix:** write it into the prompt as a final `Avoid:` line. Name the failure the
> scene invites rather than a general disclaimer — _"it does not shatter, crack,
> splinter, crumble, chip or send shards flying"_ is what bought two clean cuts
> through a glass strawberry, and a positive simile alongside it (_"parts cleanly,
> like cold butter"_) is what made the negation land.

---

## Worked examples

Three finished pieces, each with the brief that made it. Every one is a **single
generation**, and nothing here is several takes cut together.

All three came in through an image — a first frame twice and a reference once —
because that is what we have run. Pipeline A has no worked example here yet, and
a made-up one would be worth less than the gap.

### Two cuts, eight seconds

_Pipeline B: one first frame._

<figure class="guide-example guide-example--wide">
<div class="guide-example-stage"><video src="/showcase/guides/the-cut.mp4" poster="/showcase/guides/the-cut.webp" width="960" height="540" controls controlslist="nodownload noplaybackrate noremoteplayback" disablepictureinpicture playsinline preload="none" aria-label="A chef's knife cutting twice through a blown-glass strawberry on a wooden board, each cut revealing a pale cross-section"></video></div>
<figcaption class="guide-example-side">
<dl class="guide-example-spec">
<div><dt>Length</dt><dd>8s, one take</dd></div>
<div><dt>Frame</dt><dd>16:9, 1080p</dd></div>
<div><dt>Supplied</dt><dd>One first frame</dd></div>
<div><dt>Cost</dt><dd><span class="credit">7.5</span> + <span class="credit">267</span></dd></div>
</dl>
<img class="guide-example-input" src="/showcase/guides/the-cut-first-frame.webp" width="760" height="428" loading="lazy" decoding="async" alt="The supplied first frame: the whole glass strawberry on the board, knife not yet in shot" />
<p class="guide-example-input-note">The first frame, made first, to settle one question: glass, or fruit?</p>
</figcaption>
</figure>

A knife through a blown-glass strawberry — the origin format of AI ASMR, and the
one that still converts. **The sound is the product**; the picture is what makes
the sound plausible.

Two cuts in eight seconds, because one is a demo and three is a montage. Every
audio beat sits on its own `SFX:` line attached to its own timestamp, and each
one came back as a distinct, measurable event: first contact at 1.6s, the press
at 1.9s (−7.4 dB), the blade breaking through at 3.0s (−0.1 dB), the second cut
at 4.1s (−1.2 dB). **That is what the notation buys** — separately authored
sounds arriving separately, rather than a wash.

Two findings worth taking away. **Name the failure the physics invites** — glass
under a blade wants to shatter, and every early public version of this format
does. And **describe the inside before anything cuts it**, because the cut face
is the payoff shot of the entire genre and the model will not invent an interior
it was never told about.

It is also where the dead-tail rule was measured. Four beats were asked for
across eight seconds; the model played all four in the first four, and **from
4.5s every peak sits below −41 dB** — three and a half seconds, 44% of the short,
with nothing moving in the picture either. The final block asked for a ring
decaying into silence, and Veo does not fill a block that names an absence. Give
it a slice rocking to rest instead.

<details class="guide-fold"><summary>The brief, in full</summary><pre><code>Extreme close-up, one single continuous unbroken 8-second shot from a locked-off&#10;camera. No pan, no tilt, no zoom, no push-in, no dolly, no handheld drift, no&#10;cuts, no dissolves. Shallow depth of field. The framing at second 8 is identical&#10;to second 0.&#10;&#10;Subject: a solid opaque glass sculpture of a strawberry, in a strawberry's own&#10;natural red, resting on a pale warm wooden cutting board. Glossy blown-glass&#10;surface, seed pits and shoulder dimples moulded into the glass, a small opaque&#10;green glass hull on top. Inside, the glass is a paler milky pink, so every cut&#10;face reveals a clean cross-section of pale core inside red rind. It is a glass&#10;object for the whole shot and never becomes a real fruit: it never softens,&#10;never bleeds juice, never turns to wet flesh or pulp, never sweats. One human&#10;hand holds a stainless steel chef's knife and enters from the right.&#10;&#10;Context: the board sits on a dark uncluttered surface, the background falling&#10;away to soft shadow. Crisp soft studio key from camera-left, gentle specular&#10;highlights along the blade and the glass shoulders, one thin red caustic thrown&#10;onto the wood. The light, the board and the background never change.&#10;&#10;Action:&#10;&#10;[00:00-00:02] The glass strawberry sits perfectly still. The knife lowers from&#10;above and the edge settles against the top of the glass without moving it.&#10;SFX: one faint high glass tick as the steel meets the glass.&#10;&#10;[00:02-00:04] The blade presses down and passes cleanly through in one smooth&#10;unhurried stroke. The first slice separates and tips over flat onto the board,&#10;its pale cross-section facing the camera. SFX: a dense crystalline crunch&#10;building through the stroke and releasing as the blade breaks through, then a&#10;bright glassy clink as the slice lands on the wood.&#10;&#10;[00:04-00:06] The hand lifts the knife clear, shifts a finger's width to the&#10;left, and cuts again, a little faster. A second even slice falls away and comes&#10;to rest beside the first. SFX: a second crunch, shorter and sharper than the&#10;first, then a light double clink as the slice rocks twice and settles.&#10;&#10;[00:06-00:08] The knife lifts up out of the top of frame. The two slices are&#10;still. Nothing else moves. SFX: the last of the glass ringing decaying away into&#10;silence.&#10;&#10;Cutting behaviour: the glass parts cleanly, like cold butter, along the exact&#10;line of the blade. It does not shatter, crack, splinter, crumble, chip or send&#10;shards flying, and no piece leaves the board.&#10;&#10;Ambient noise: a quiet padded studio room tone underneath everything, and&#10;nothing else.&#10;&#10;Style: photorealistic, ultra sharp, calm and premium, high-fidelity ASMR&#10;recording in both picture and sound. The audio is close-miked and intimate, and&#10;the loudest thing in the mix is the glass.&#10;&#10;Avoid: music, score, melody, humming, singing, speech, voice-over, narration,&#10;breathing, a second hand, a second piece of fruit, real fruit juice or pulp,&#10;flying shards, camera movement of any kind, cuts, crossfades, background&#10;changes, extra props, on-screen text, captions, subtitles, watermarks and logos.</code></pre></details>

### A commute nobody filmed

_Pipeline D: one reference._

<figure class="guide-example guide-example--wide">
<div class="guide-example-stage"><video src="/showcase/guides/the-commute.mp4" poster="/showcase/guides/the-commute.webp" width="960" height="540" controls controlslist="nodownload noplaybackrate noremoteplayback" disablepictureinpicture playsinline preload="none" aria-label="A yeti films himself on a snow ridge above a sea of cloud, talks about his commute, is hit by a gust of spindrift and wipes the lens clear with one finger"></video></div>
<figcaption class="guide-example-side">
<dl class="guide-example-spec">
<div><dt>Length</dt><dd>8s, one take</dd></div>
<div><dt>Frame</dt><dd>16:9, 1080p</dd></div>
<div><dt>Supplied</dt><dd>One reference image</dd></div>
<div><dt>Cost</dt><dd><span class="credit">7.5</span> + <span class="credit">267</span></dd></div>
</dl>
<img class="guide-example-input" src="/showcase/guides/the-commute-reference.webp" width="760" height="428" loading="lazy" decoding="async" alt="The supplied reference: the yeti standing on the ridge, one arm already reaching out of frame toward the lens" />
<p class="guide-example-input-note">The one reference — and the arm reaching out of frame is the rig, already built.</p>
</figcaption>
</figure>

The cryptid vlog is where character AI video started, and the joke is that he
means it sincerely. Eight seconds of a chore, filmed by the person doing it: he
walks, he says two sentences about the walk, the weather interrupts him, and he
finishes the sentence.

**The beat sheet was obeyed almost exactly.** Transcribed with word timestamps
against the four blocks asked for:

| Asked           | Got              | Line                         |
| --------------- | ---------------- | ---------------------------- |
| `[00:00-00:02]` | 0.00–1.60        | "Forty minutes to the ridge" |
| `[00:02-00:04]` | 1.64–4.10        | "Forty minutes back"         |
| `[00:04-00:06]` | 4.10–6.20 silent | the gust                     |
| `[00:06-00:08]` | 6.20–7.38        | "Every day of my life"       |

Four for four — and the strongest evidence it read the sheet is the **2.1 seconds
of deliberate silence** in the gust window. It held its tongue across a block
that named an action and no dialogue, then came back in on time.

The wipe rendered, and in the right order: the gust whites the frame at ~4.0s,
the finger enters from the right at ~5.0s _while it is still white_, and the
picture clears behind it. Three ideas held at once — the lens is a surface, snow
is on it, and a finger belongs to the hand not carrying the stick.

**And he does not walk.** The brief asks for it in the shot line, in the subject
paragraph and in every one of the four blocks. The reference is a standing
portrait, and standing is what came back — the whole of the pose rule, in one
clip.

<details class="guide-fold"><summary>The brief, in full</summary><pre><code>A single continuous unbroken 8-second handheld selfie video, filmed by the yeti&#10;himself. He is holding a selfie stick in his right hand, at arm's length, a&#10;little above his own eyeline and angled down at him — that is where the camera&#10;is. The camera never leaves his hand. Wide-angle lens, his head and shoulders&#10;filling the frame, the ridge and the sky behind him. The frame jolts and swings&#10;with every step he takes, and it is already moving at second 0.&#10;&#10;Subject: the creature from the reference image, unchanged — the same dirty&#10;off-white shaggy fur, the same broad flat face, the same tired patient&#10;expression. He is walking for the whole shot and never stops, never sits and&#10;never puts the camera down. He speaks directly into the lens in a deep, low,&#10;rumbling voice, flat and matter-of-fact, like a man describing his commute. His&#10;breath fogs in the cold.&#10;&#10;Context: he is trudging along an exposed snow ridge high above a sea of cloud,&#10;exactly as in the reference image, the wind hard from his left and spindrift&#10;streaming past him. Bright flat diffuse daylight with no sun disc, no cast&#10;shadows and no rim light. The light never changes and the background stays a&#10;soft white cloud layer throughout.&#10;&#10;Action:&#10;&#10;[00:00-00:02] Already mid-stride and already shaking. He glances into the lens&#10;and says, "Forty minutes to the ridge."&#10;SFX: the deep crunch and squeak of huge feet breaking wind-packed snow, one&#10;heavy step at a time, close to the microphone.&#10;&#10;[00:02-00:04] He keeps walking, his breath fogging past the lens, and says,&#10;"Forty minutes back." The wind rises behind him.&#10;SFX: fur and wind buffeting the microphone, a low gust building.&#10;&#10;[00:04-00:06] A hard gust hits from his left and throws a sheet of spindrift&#10;straight across the lens. The picture goes white and grainy and he is only a&#10;dark shape behind it. He does not react and does not stop walking.&#10;SFX: a sharp roar of wind peaking, snow hissing against the lens.&#10;&#10;[00:06-00:08] One enormous furred finger comes into frame from the right and&#10;drags across the lens, smearing the snow clear in a wet streak with a little&#10;still caught in the corners. His face comes back into view, still walking, and&#10;he finishes, "Every day of my life."&#10;SFX: a soft muffled scrape of fur on wet glass, then the footfalls again.&#10;&#10;Ambient noise: a constant high-altitude wind under everything, thin and cold,&#10;with no melody in it.&#10;&#10;Style: photorealistic amateur vlog footage, natural available light only — the&#10;honest look of a phone camera in bad weather, with mild wide-angle distortion, a&#10;little sensor noise, and an exposure that hunts slightly in all the white.&#10;&#10;Avoid: music, score, soundtrack, humming, singing, a narrator, a second voice, a&#10;second character, any other person or animal, on-screen text, captions,&#10;subtitles, titles, watermarks, logos, a locked-off or tripod-steady frame, a&#10;camera that leaves his hand, a third-person shot, a drone shot, cuts, dissolves,&#10;speed ramps, slow motion, the yeti stopping or standing still, bared fangs, a&#10;snarl, a costume seam, a human face.</code></pre></details>

### The face Seedance would not take

_Pipeline B on Veo 3.1 Fast: one first frame, vertical._

<figure class="guide-example">
<div class="guide-example-stage"><video src="/showcase/guides/the-ad.mp4" poster="/showcase/guides/the-ad.webp" width="608" height="1080" controls controlslist="nodownload noplaybackrate noremoteplayback" disablepictureinpicture playsinline preload="none" aria-label="A woman in a parked car holds a mint carton up beside her face and talks through it in four short lines, turning the box and shaking a sachet"></video></div>
<figcaption class="guide-example-side">
<dl class="guide-example-spec">
<div><dt>Length</dt><dd>8s, one take</dd></div>
<div><dt>Frame</dt><dd>9:16, 720p</dd></div>
<div><dt>Supplied</dt><dd>One first frame</dd></div>
<div><dt>Cost</dt><dd><span class="credit">11</span> + <span class="credit">100</span></dd></div>
</dl>
<img class="guide-example-input" src="/showcase/guides/the-ad-first-frame.webp" width="540" height="960" loading="lazy" decoding="async" alt="The supplied first frame: the same woman in the driver's seat holding the mint carton beside her face" />
<p class="guide-example-input-note">The first frame, made first, so the brand on the carton was settled before anything moved.</p>
</figcaption>
</figure>

A UGC ad in the shape almost every one of them takes: four beats, four quoted
lines, a fixed camera on the dashboard, and a product held up beside a face.

**It exists because Seedance refused it.** The identical first frame came back
`E005 — the input or output was flagged as sensitive`, which is the category that
filter is strictest about, and re-rolling an identical input against the same
classifier is not a plan. Veo 3.1 Fast took the same picture without comment.

Two smaller things it demonstrates. **A first frame keeps 9:16** — the output is
720×1280, and this is the door a vertical piece has to come through, because
references would have forced it to landscape. And **four quoted lines fit eight
seconds** with room to breathe: eighteen words in total, under the twenty-word
ceiling, and none of them ends mid-syllable.

The one thing it does not demonstrate is small print. Watch the carton across the
eight seconds — the mint and the shape hold, but the wordmark under them gives up
around the halfway mark. **Large type survives a video model. Small type is
texture**, and text that has to stay readable belongs in the first frame, large
in frame, not in the prompt.

<details class="guide-fold"><summary>The brief, in full</summary><pre><code>One continuous unbroken 8-second vertical shot from a completely fixed camera&#10;resting on the dashboard. No pan, no tilt, no zoom, no push-in, no handheld&#10;drift, no cut, no dissolve. The framing at second 8 is identical to second 0.&#10;&#10;Subject: the woman of the first frame, unchanged in every frame — same face,&#10;same hair, same dusty-pink fleece, same seatbelt, same car. She stays in the&#10;driver's seat of the parked car throughout and never opens a door, never starts&#10;the engine, never turns away from the lens.&#10;&#10;Product: the pale mint NUVELO carton of the first frame, its lettering unchanged&#10;and legible whenever it faces the camera, plus one loose stick sachet in the&#10;same mint and coral. The carton never changes size, colour or wording.&#10;&#10;Action:&#10;&#10;[00:00-00:02] She is already holding the carton up beside her face. She looks&#10;into the lens and speaks, her other hand resting on the wheel.&#10;"Okay, I have to talk about these."&#10;SFX: the close soft ambience of a parked car interior.&#10;&#10;[00:02-00:04] She rotates the carton a quarter turn to show its side, then&#10;brings it square to the camera again.&#10;"Thirty sticks, one a day."&#10;SFX: a light card-on-card rustle as the box turns in her hand.&#10;&#10;[00:04-00:06] She lowers the carton out of the bottom of the frame and brings up&#10;a single stick sachet, pinched between thumb and forefinger, and gives it one&#10;small shake.&#10;"You just tear one open."&#10;SFX: a short crisp foil crinkle as the sachet shakes.&#10;&#10;[00:06-00:08] She raises the carton back up beside her face so the box and the&#10;sachet are held up together, one in each hand, nods once and smiles.&#10;"That's the whole thing. Try it."&#10;SFX: nothing but the room.&#10;&#10;Light: constant flat daylight through the windshield for all eight seconds. No&#10;exposure shift, no relight, no flicker, no colour drift.&#10;&#10;Ambient noise: the quiet interior of a parked car, a faint distant street&#10;outside, and nothing else.&#10;&#10;Style: plain front-facing phone video, no colour grade, no cinematic contrast,&#10;faint sensor noise in the shadows. Her voice is close, warm, conversational and&#10;unhurried, as recorded by the phone's own microphone.&#10;&#10;Avoid: music, score, backing track, a second voice, a second person, a second&#10;pair of hands, camera movement of any kind, cuts, crossfades, speed ramps, the&#10;car moving, a face or hair that morphs or re-rolls, clothes that change, the&#10;carton changing its wording or colour, a phone visible in frame, on-screen text,&#10;captions, subtitles, watermarks and logos.</code></pre></details>

---

## About justpictur.ing

justpictur.ing is an AI picture studio: prompt-driven images, clips, voice-overs and music, and storyboard-backed films with consistent characters across every scene.

Agents connect over MCP and can generate directly:

```
claude mcp add justpicturing https://api.justpictur.ing/mcp --transport http
```

The app is at https://app.justpictur.ing and requires an account.

Every guide on this site has a Markdown twin at its own URL plus `.md`. Every guide: https://justpictur.ing/guides/
