---
title: "P-Video: the complete guide to Pruna's four models"
description: "Four models, one family: a cheap generator and three performance rigs. What each one takes, what a take really costs, and the failures worth knowing first."
canonical_url: "https://justpictur.ing/guides/p-video-prompting/"
published: 2026-08-15
updated: 2026-08-16
locale: "en"
models: ["p-video", "p-video-avatar", "p-video-animate", "p-video-replace", "gpt-image-2", "gemini-3.1-flash-tts"]
---

Pruna makes four video models, and they do not compete with each other. One
makes a clip out of a prompt. The other three animate something you already
have: a portrait and a voice, a still and a performance, a shot and a new face
in it.

**They have become the models we reach for first**, and the reason is
arithmetic. Twenty seconds of someone talking to camera costs
<span class="credit">31.5</span> here and <span class="credit">200</span> on the
model above it. Five seconds of 720p costs <span class="credit">8.5</span>,
against <span class="credit">37.5</span> for the next cheapest model that will
render the same shot. What comes back is good enough that the saving stops
being a compromise and starts being the default.

> Four models is not four things to learn. It is one question — _what is already
> in your hands?_ — asked four ways.

**NOTE:** Everything here is true of all four unless it names one.

## Overview

### The short version

Four decisions do most of the work. If you read nothing else, read these. Each
one is written out in full further down.

<ol class="guide-keys">
<li class="guide-key">
<p class="guide-key-title">Pick the model from what you already have, not from what you want</p>
<p class="guide-key-body">The four split cleanly by input and there is no overlap between them. <strong>Three of them refuse to start without media you supply</strong>, so the model is chosen for you the moment you know whether you are holding a voice, a video, or nothing.</p>
<div class="guide-key-lines">
<p class="guide-key-practice">Nothing but words → P-Video. A portrait and a voice → Avatar. A still and a performance to copy → Animate. A shot whose person should be somebody else → Replace.</p>
</div>
</li>
<li class="guide-key">
<p class="guide-key-title">A performance take is a flat price, whatever comes back</p>
<p class="guide-key-body">Avatar, Animate and Replace read their length off the media you attached, so the take is quoted at the fifteen-second ceiling and never above it. <strong>A 4.5-second clip and a 17.9-second one came to the same number</strong> — so a short line wastes the budget a long one would have used for free.</p>
<div class="guide-key-lines">
<p class="guide-key-practice">Write to the ceiling. Two sentences and six sentences cost the same, so put the whole thought in one take rather than splitting it across two.</p>
</div>
</li>
<li class="guide-key">
<p class="guide-key-title">What you attach outranks what you write</p>
<p class="guide-key-body">On the three performance models the driver carries the timing, the motion and the performance, and the prompt only decorates what is left. <strong>A property the media sets cannot be argued out of the model in prose</strong> — asked for "curious, then a chuckle, then a complicated smile", the model returned one continuous pleasant smile, because it drives the mouth from the waveform.</p>
<div class="guide-key-lines">
<div class="guide-key-quote guide-key-quote--no"><p class="guide-key-quote-label">Instead of</p><p class="guide-key-quote-text">Writing the expression you want into the prompt</p></div>
<div class="guide-key-quote guide-key-quote--yes"><p class="guide-key-quote-label">Do</p><p class="guide-key-quote-text">Put it in the still you seed with, or in the take you drive with</p></div>
</div>
</li>
<li class="guide-key">
<p class="guide-key-title">P-Video has no reference slot — it has frames</p>
<p class="guide-key-body">You cannot hand it a character and ask for that character. What you can hand it is the exact picture the clip opens on, the exact picture it arrives at, or both. <strong>A supplied frame also overrules your <span class="term"><button class="term-word" type="button" popovertarget="def-aspect-ratio">aspect ratio</button><span class="term-def" popover id="def-aspect-ratio">The shape of the frame: its width against its height. 16:9 is the wide shape a television is; 9:16 is a phone held upright.</span></span> setting</strong>, because the ratio comes off the image.</p>
<div class="guide-key-lines">
<p class="guide-key-practice">Settle the subject in one image first, then hand that image over as the Start frame. Consistency comes from the picture here, never from the prompt.</p>
</div>
</li>
</ol>

### The family, in one table

| Model               | Needs                                 | Also takes                              | Length                | Frame shape                   | Sound                    |
| ------------------- | ------------------------------------- | --------------------------------------- | --------------------- | ----------------------------- | ------------------------ |
| **P-Video**         | a prompt                              | Start frame, End frame, one audio track | 1–10s, or the audio's | seven, unless a frame sets it | generated, switchable    |
| **P-Video Avatar**  | a portrait **and** a voice track      | an optional prompt                      | the voice's, exactly  | the portrait's                | the voice you gave it    |
| **P-Video Animate** | an image **and** a source video       | an optional prompt                      | the source's          | the source's                  | the source's, always on  |
| **P-Video Replace** | a source video **and** 1–3 identities | an optional prompt                      | the source's          | the source's                  | the source's, switchable |

All four render at **720p or 1080p**, and all four take a prompt of up to
**4,000 characters** — twice what an image model gets, and more than any of them
needs. On three of the four the prompt is optional, and one of the worked
examples below ran with it empty.

### What it costs

**P-Video is priced by the second**, and it is the cheapest way in the catalog
to turn a prompt into a moving picture — $0.02 a second at 720p, under half the
rate of anything else that renders at that size.

| Length | 720p                            | 1080p                            | What it's for                                             |
| ------ | ------------------------------- | -------------------------------- | --------------------------------------------------------- |
| 1s     | <span class="credit">2</span>   | <span class="credit">3.5</span>  | A punctuating cut. One second is a real option here       |
| 3s     | <span class="credit">5</span>   | <span class="credit">10</span>   | A reaction, a beat, a shot with one thing happening in it |
| 5s     | <span class="credit">8.5</span> | <span class="credit">17</span>   | The default. A complete idea                              |
| 10s    | <span class="credit">17</span>  | <span class="credit">33.5</span> | The ceiling, and long enough to hold two beats            |

**The other three are a flat price.** The length belongs to the media you
attached, so no duration is ever sent and the take is quoted at the ceiling —
fifteen seconds' worth — with settlement never going above it. Every take we
have run has come to exactly this number, whether the clip was four seconds or
twenty:

| Model               | 720p                             | 1080p                            |
| ------------------- | -------------------------------- | -------------------------------- |
| **P-Video Avatar**  | <span class="credit">31.5</span> | <span class="credit">56.5</span> |
| **P-Video Animate** | <span class="credit">37.5</span> | <span class="credit">75</span>   |
| **P-Video Replace** | <span class="credit">37.5</span> | <span class="credit">75</span>   |

<blockquote class="guide-quote--caution">
<p><strong>Fifteen seconds is where the price stops, not where the clip stops.</strong> The composer shows the ceiling because it is what gets held, and the model never receives a duration at all — a 17.9-second line and a 20.16-second one both came back whole, and both cost <span class="credit">31.5</span>.</p>
<p>So do not split a part to fit a number that is not a limit. Split it at a real seam, or not at all.</p>
</blockquote>

### Why they became the default

The same three jobs, priced against what else in the catalog can do them:

| The job                                  | Pruna                                                | What else does it                                                                                                                                                                                         |
| ---------------------------------------- | ---------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Five seconds of 720p clip                | **P-Video** <span class="credit">8.5</span>          | Seedance 2.0 Mini <span class="credit">37.5</span> · [Seedance 2.5](/guides/seedance-2-5-prompting/) at 480p <span class="credit">43</span> · Runway Gen-4.5 <span class="credit">50</span>               |
| Twenty seconds of someone talking        | **P-Video Avatar** <span class="credit">31.5</span>  | OmniHuman 1.5 <span class="credit">200</span> · three 8s [Veo 3.1 Fast](/guides/veo-3-1-prompting/) takes joined <span class="credit">300</span> · Seedance 2.5 at 720p <span class="credit">385.5</span> |
| Recasting a performance onto a character | **P-Video Animate** <span class="credit">37.5</span> | Kling Motion Control <span class="credit">87.5</span>                                                                                                                                                     |

<blockquote class="guide-quote--caution">
<p>These are not the best models in the catalog, and the guide does not claim they are. What they are is the models where <strong>the extra spend stops buying anything a viewer notices</strong> — and that is a different, more useful property.</p>
<p>Reach past them when the face is the performance rather than the thing the words come out of, when the camera has to do something cinematic, or when a <em>generated</em> shot needs to run longer than ten seconds in one piece.</p>
</blockquote>

---

## The rules

Seven things that are true whatever you are making. Each one cost us money to
learn, and they are ordered by how much they will save you.

### Choose from your inputs, not from your ambition

> Three of the four refuse to start. What they refuse over is the model choice,
> made for you.

A generation here is blocked until its required seats are full, and that is the
whole taxonomy:

- **Avatar** wants an image and an audio track.
- **Animate** wants an image and a video.
- **Replace** wants a video and between one and three identities.
- **P-Video** wants nothing at all.

We hit this on a reaction shot: three seconds of a woman watching something,
saying nothing. No line meant no audio, which meant no length and no performance
for a lip-sync model to build from; there was no footage to recast either.
**P-Video was the only one of the four that could take the shot**, and it took
it for <span class="credit">13.5</span> at 1080p.

Reading the requirement backwards is the fast route to the right model. If you
are reaching for Avatar and have no voice take yet, the job is not "generate a
clip", it is "cut the voice first".

### A performance take is a flat price, so write to the ceiling

> The clip is as long as what drives it, and the bill is as long as fifteen
> seconds either way.

None of the three performance models takes a length. Every request is the auto
sentinel, no number ever reaches the provider, and the take is quoted at the
fifteen-second ceiling regardless of what the media will turn out to be worth.
Measured across our own runs:

- A **4.48-second** avatar clip and a **17.9-second** one both came to
  <span class="credit">31.5</span> at 720p.
- A twenty-second one, in the [worked example](#a-host-who-does-not-exist)
  below, came to the same <span class="credit">31.5</span>.

That inverts the usual instinct. On a per-second model you write short to save
money; here writing short **wastes** the seconds you already paid for. Six
sentences with two marked pauses is twenty seconds and one take. The same
content split across two takes is twice the price and a cut you have to hide.

The corollary is the budgeting rule: **a lip-synced piece is budgeted from the
script, not from the clip.** The only way to make a shorter take is to write a
shorter line.

### The driver is the clock, so cut it first

> The audio decides the length, which means the audio has to exist before the
> picture does.

On Avatar this is exact rather than approximate. A 967,724-byte 24kHz mono wav
is 20.16 seconds and the clip came back 20.16 seconds. Across four more takes
each clip landed **within 50 milliseconds** of the voice-over that drove it, and
the delivered soundtrack correlated with the source mp3 at **0.9997** — the
model re-emits the track you gave it rather than resynthesising anything.

So the running order is voice, then picture, in that order, every time. Do it
the other way round and you are trimming a face to fit a line.

The same rule holds one step sideways on Animate and Replace, where the source
video is the clock: its length, its timing, its camera and its frame rate all
survive into the output, and the only way to change any of them is to change the
source.

### What the picture sets, the prompt cannot argue out

> A property the input carries is not up for discussion downstream.

This is the single most expensive thing to learn about the family, and it shows
up in three different places.

**On expression.** Given _"curious, then a chuckle, then a complicated smile"_,
Avatar returns one continuous pleasant smile. It drives the mouth from the
waveform and ignores the prompt's semantics. When the face _is_ the performance
and the audio is nearly silent, that is what OmniHuman 1.5 costs
<span class="credit">200</span> for.

**On gesture.** A hand in the seed still is a beat the model returns to, not a
pose it holds. On a ten-second take a raised index finger was up at 0.2s and
1.5s, gone by 3s, **back at 7s**, gone by 9s. On takes under about six seconds
the hands play once at the top and settle out of frame. So a gesture still buys
you a reliable **opening beat** and nothing you can place — never write a shot
that needs a hand present at a chosen moment.

**On framing.** The output takes its shape from the portrait or from the source
video, not from the ratio you selected. A 16:9 request came back **1280×704**,
and a vertical take came back **1088×1920** — eight pixels of encoder padding on
a 1080-wide composition. The setting shapes the composer's preview; the file is
the authority. Crop, do not scale, or every measured coordinate in your edit
moves.

### Build from stills, not from a chain

> Each hop re-renders the last one, and the drift is monotonic.

It is tempting to run a sequence by feeding each clip's final frame into the
next: the boundaries carry no visible cut, and it is one fewer image to make. We
built a four-part sequence that way and measured face-region difference against
the original anchor at each step:

**29.2 → 35.2 → 36.9 → 41.2.**

A clean climb, which is what generational drift looks like. Rebuilt from a fresh
still per part, the same measure reads **35.6 · 40.2 · 43.6 · 32.7** —
unordered, because what it now tracks is pose rather than accumulation. Chaining
also freezes one pose for the whole piece: she never gestures, never sits back,
never does anything but talk from where the previous clip left her.

> **Chain only where the performance genuinely continues.** One boundary in that
> sequence earned it — she is still smiling at the screen when the next line
> starts, and she lifts her eyes on it. Frame 0 of the second clip differed from
> the first clip's final frame by **3.78/255**, against a bare-wall patch that
> could not have changed at **3.99**. Same magnitude: what separates them is
> re-encode noise and nothing else.

Everywhere else, a fresh still is a hard cut and the only way a shot gets its
own pose. Cheap, too — one image is a fraction of the clip it seeds.

**The drift has a second signature, and this one you can see.** Face difference
is a number; colour is a picture. Mean chroma across a fixed crop of the face
falls inside every take that runs long enough, and a continuation opens on the
value the previous take closed at — so the fall carries across the boundary
instead of resetting at it. Two takes from one session, both 9:16, one identity,
four minutes apart:

<figure class="guide-example guide-example--wide">
<img class="guide-example-plate" src="/showcase/guides/the-chain-two-takes.webp" width="1595" height="603" loading="lazy" decoding="async" alt="Two rows of six face crops of the same woman. The top row is warm-skinned and freckled; the bottom row is grey, waxy and hollow-cheeked." />
<figcaption class="guide-example-side"><p class="guide-example-input-note">Top: a take seeded from its own still, frames 80–110. Bottom: a take that opened on the previous clip's final frame, frames 84–114. Same wardrobe, same wall, same light, four minutes apart.</p></figcaption>
</figure>

The top take loses **8.6%** of its face colour over 4.7 seconds and is still
recognisably her at the end of it. The bottom take is one hop further on. It
opens **28% below** where the top one opens, closes **44% below**, and spends
22.6% inside its own 4.9 seconds — the steepest of the nine takes we measured
that night, and the only one to finish below 60% of where the top take opened.
What that number looks like is the picture: pallor, waxy skin with the pore
detail gone, hollow cheeks, sunken sockets, an older face.

**The handoff is literal.** Frame 0 of the failing take and the final frame of
the take before it are not similar, they are the same picture — 36.9 dB PSNR,
against 16.7 dB for an unrelated pair from the same session.

<figure class="guide-example guide-example--wide">
<img class="guide-example-plate" src="/showcase/guides/the-chain-handoff.webp" width="1550" height="449" loading="lazy" decoding="async" alt="Four face crops: a take's first and last frame, then the next take's first and last frame. The middle two are identical; the fourth is grey and drawn." />
<figcaption class="guide-example-side"><p class="guide-example-input-note">The parent take's first and last frame, then the child's. The middle two are one picture, handed over. The fourth is 4.9 seconds after it.</p></figcaption>
</figure>

**The still sets the budget, and the take spends it.** A first-hop take opens
within 4% of the colour of the still it was seeded from — in all four pairs we
could match, the model preserved what it was given and then bled it. So a chain
does not go wrong at the end. It goes wrong at the picture you hand it, and
every hop after that spends what is left:

| Take | Opened on                | Length | Opens | Closes  |
| ---- | ------------------------ | ------ | ----- | ------- |
| 1    | its own still            | 3.0s   | 105%  | 99%     |
| 2    | its own still            | 3.0s   | 101%  | 103%    |
| 3    | its own still            | 2.9s   | 95%   | 96%     |
| 4    | its own still            | 5.5s   | 103%  | 87%     |
| 5    | its own still            | 6.7s   | 92%   | 75%     |
| 6    | its own still            | 4.7s   | 100%  | 91%     |
| 7    | take 6's final frame     | 10.6s  | 94%   | 86%     |
| 8    | its own still            | 10.6s  | 79%   | 71%     |
| 9    | **take 8's final frame** | 4.9s   | 72%   | **56%** |

Nine takes, one session, one identity, one room. Each figure is mean
face-region chroma over the take's first and last tenth, against take 6's
opening. Two things spend the budget and they are not the same thing:
**length** — three seconds costs almost nothing, ten costs about a tenth — and
**hops**. Take 7 chains and survives it, because it opened at 94%. Take 9 chains
from a take that had already opened low, and lands somewhere no prompt and no
re-roll recovers.

**Which makes it cheap to catch.** One pass of ffmpeg's `signalstats` over a
crop of the face reports chroma per frame, and two thresholds separate the good
takes from the corpse here without anything cleverer: a take that closes more
than about 15% below where it opened is drifting, and a take that opens below
the still it was seeded from has inherited a used frame.

### It will animate a photographed face

> The refusal you hit elsewhere is not here, and for a lot of work that is the
> whole reason to be here.

Hand [Seedance 2.5](/guides/seedance-2-5-prompting/) a photorealistic close-up
of a person and it declines — seven refusals in seven attempts in our tests,
through every door. Pruna's models take the same picture without comment.

That is the axis on which the choice actually turns for
<span class="term"><button class="term-word" type="button" popovertarget="def-ugc">UGC</button><span class="term-def" popover id="def-ugc">User-generated content: an ad shot to look like a real customer filmed it on their phone, rather than like an ad.</span></span>
work: if your piece needs a believable person talking to camera from a picture
you already have, this family will do it. So will
[Veo 3.1](/guides/veo-3-1-prompting/), which is the stronger speech model — but
Veo stops at eight seconds and charges <span class="credit">100</span> to get
there on its Fast tier, where twenty seconds here is
<span class="credit">31.5</span>.

**It is not a licence.** A face you have no right to use is still a face you
have no right to use. Every person in every example on this page was generated,
and nothing was photographed.

### Test at 720p, and let the price do the testing

> 1080p costs double on three of the four, and a little under double on Avatar.
> Nothing else about a take changes.

There is no rehearsal model here and no draft tier — the provider has one, and
the catalog deliberately does not wire it up, because the full-quality run is
already cheaper than most models' rehearsals. Five seconds of P-Video at 720p is
<span class="credit">8.5</span>; the same brief at 1080p is
<span class="credit">17</span>. Dropping a tier changes sharpness and changes
nothing a prompt controls.

So the working method is: **run it at 720p and read the result honestly.** It
answers everything a brief can get wrong — whether the framing is what you
pictured, whether the motion holds, whether your inputs are accepted, whether
the character survives. Then re-run the same words at 1080p for the version you
keep.

One provider note that cuts the other way: vertical work is reported to hold up
better at 1080p. If a 9:16 take looks soft in a way 720p usually does not
explain, that is the first thing to try.

---

## The four models

Each one is a different job. They are written here in the order you meet them:
the one that needs nothing, then the three that need something.

### P-Video — the one that needs nothing

#### What it is for

- **Best for:** <span class="term"><button class="term-word" type="button" popovertarget="def-b-roll">b-roll</button><span class="term-def" popover id="def-b-roll">Footage with nobody speaking in it, like scenery, hands, a street or a detail, cut in around the main shots to cover a voice-over or to let a scene breathe.</span></span>, close-ups, product animation, reaction shots with no line in them, and the connective tissue of a longer piece.
- **Costs you:** reference images. There is no slot for one, so a subject cannot be carried in.
- **Watch for:** the aspect ratio comes from a supplied frame, not from your setting.

#### What it takes

Seven aspect ratios, one to ten seconds, 720p or 1080p, and three optional
inputs that change what the model is doing:

```
Start frame     the clip opens on this exact picture
End frame       the clip arrives at this exact picture
Audio track     the clip is cut to this, and takes its length from it
```

Attach an audio track and the duration control goes away: the sound decides,
exactly as it does on the performance models. That is the cheapest music-video
route in the catalog, and it is unusual to find an audio input at this price at
all.

**Sound is generated by default and costs nothing extra.** Unlike
[Veo 3.1](/guides/veo-3-1-prompting/), where the audio switch doubles the rate,
speech and effects here come out of the same per-second price whether you write
for them or not. The provider is candid that sound effects are the weaker half —
for anything where the sound is the product, make the track in the Audios room
and hand it back as the audio input.

#### How to write it

The provider's own formula is a good starting shape, and only the first three
parts are required:

```
[Subject] [Action] [Scene] [Camera] [Lighting] [Style] [Audio]
```

Two habits carry over from the rest of this shelf and matter more than the
formula does. **Name what is visible in the frame rather than the camera term
for it** — a description of what should be on screen gets followed where an
angle gets interpreted. And **describe anyone who has to stay the same once**,
in detail, then point back to that; describing them twice invites a redraw.

The one thing the prompt cannot do is hold a subject across takes. Nothing is
carrying identity, so if the same person has to appear in two clips, settle them
in an image first and hand that image in as the Start frame.

#### The known limits

The model's own README is unusually direct about where it stops, and our runs
agree: **no extreme cinematic camera motion, no native 4K, weak sound effects,
and speaker attribution that drifts above two voices.** It is very good at
close-up subjects and foreground objects, and it is not a model to plan a
multi-scene story around.

**Worked example:** [Three seconds of nothing happening](#three-seconds-of-nothing-happening)
— a reaction shot no other model in the family could take.

### P-Video Avatar — a portrait and a voice

#### What it is for

- **Best for:** anything scripted to camera — a host, an explainer, a UGC ad, a localised version of a piece you already made.
- **Costs you:** the choice of length, and any control over the performance beyond what the still and the track already carry.
- **Watch for:** it drives the mouth from the waveform, so a silent moment is not a performance.

#### What it takes

Two required seats — **Avatar image** and **Voice audio** — and an optional
prompt. That is the whole recipe. The audio's voice and timing animate your
image, and the clip is as long as the track — to the frame, and past fifteen
seconds without paying for it.

The provider will also synthesise speech itself, from a script and one of thirty
named voices. **We deliberately do not wire that up.** The Audios room already
writes voice-overs across five TTS models, so a take from there is the avatar's
voice, and there is one voice picker in the product rather than two. It also
makes the lineage legible: the voice exists as its own generation, with its own
prompt and its own cost, and the script can be rewritten and re-cut without
touching the host at all.

#### How to write it

> This is the model where **less prompt is better**, and often none at all is
> best.

The two required inputs already specify almost everything: who is talking and
what they are saying. The prompt is visual direction on top of that, and the
portrait has usually settled the visual direction already.

- **Use a clear, front-facing portrait.** Heavy angles, occlusion and low
  resolution all cost you identity.
- **Keep the audio clean.** Clear speech with little background noise gives
  tighter sync.
- **Put the performance in the still**, not in the prompt — see
  [the rule above](#what-the-picture-sets-the-prompt-cannot-argue-out).

#### The known limit

**It will not play a scripted expression arc.** The mouth follows the waveform
and the semantic content of the prompt does not reach the face. It is right when
the line carries the performance; when the face carries it and the audio is
nearly silent, this is not the model.

One more, and it is the kind you only catch by looking: **the face drifts on
some samples, and the still is not at fault.** Two clips in one run came back
narrower and harder with patchier skin while all six source stills measured
identical. A re-roll fixed it. Check the delivered clips, not the images.

**Worked example:** [A host who does not exist](#a-host-who-does-not-exist):
twenty seconds, two inputs, and an empty prompt.

### P-Video Animate — a performance, recast

#### What it is for

- **Best for:** putting a character you designed into a movement you already have — a walk, a dance, a piece of found choreography, a performance you filmed once and want to reuse.
- **Costs you:** any control the source video already exercises. The pace, the camera, the timing and the sound all arrive with it.
- **Watch for:** the image is **identity, not frame zero**. It is not a Start frame and it will not be treated as one.

#### What it takes

**Image to animate** plus **Driving performance video**, and an optional prompt.
The subject of your image performs the motion of the source, and the source's
audio is muxed onto the output — always. There is no switch, which is why the
model reads as sound-required in the composer.

That is the difference between this and generating the same shot from a prompt.
A motion transfer proves you can move a performance onto somebody else; a
generated version proves the model can invent a walk. Both are true statements
about the product, and only one of them needs footage.

#### How to write it

> The prompt is direction **on top of** a performance. It is not the performance.

This is why an Animate brief is four sentences where the equivalent
prompt-only brief runs four paragraphs: everything a prompt would otherwise have
to choreograph — the pace, the arm that stays in the pocket, the coat hem, the
gaze arriving on the last two steps — is already in the source and does not need
writing down.

What the source cannot supply is what to preserve during the transfer, and that
is what the prompt is for: performance emphasis, and what must remain unchanged.

- **Strong, well-lit motion transfers.** Subtle or occluded movement does not,
  and a busy handheld source is the usual reason a transfer fails.
- **A front-facing reference image** preserves identity; heavy angles do not.
- **Frame the source the way you want the result framed**, because the output
  keeps the source's aspect ratio and its frame rate.

#### Where to start

The catalog has a published recipe on this model — **Fashion Walk / Outfit
Reel** — which takes a styled-look image and a walking performance and gives you
the two seats already labelled. It is the fastest way to see the shape of the
job before you build one of your own.

### P-Video Replace — the same shot, a different person

#### What it is for

- **Best for:** keeping a shot and changing who is in it — ad variants across markets, a mascot dropped into existing footage, one strong performance reused with different talent.
- **Costs you:** the scene. This model does not invent one; it preserves the one you gave it.
- **Watch for:** the result is only as good as the match between your subject's proportions and the original's.

#### What it takes

**Source video** plus one to three **Replacement identities**, and an optional
prompt. Motion, timing, camera and scene all stay exactly as they were; the
person changes. The source's frame rate is pinned to `original`, so nothing is
resampled behind your back.

It is the only one of the three performance models where **the sound is still a
choice** — Animate welds the source track on, and this one leaves the switch
live.

#### How to write it

Three identity slots rather than one is the detail that matters, and it is worth
spending them. Three genuinely different views of the same subject is what makes
a replacement hold across a shot with camera movement in it; three frames of the
same pose is one reference with extra steps.

Otherwise the guidance is Animate's, for the same reasons: **clear, front-facing
references with minimal occlusion**, and **strong, well-lit motion** in the
source. The prompt's job is to say how the people from the references should sit
in the scene, and nothing more.

### Putting them together

The four are worth more as a chain than as four models, and the chain is short:

```
an image model      →  a still that settles who and where
P-Video             →  the shots with nobody speaking
P-Video Avatar      →  the shots with a line, one still and one voice take each
P-Video Animate     →  a shot you already have a performance for
P-Video Replace     →  a shot you already have entirely
```

A twelve-part vertical short we built runs almost exactly that way: one anchor
image holding the character, a fresh still per part changing only the eyeline
and the hands, one P-Video clip for the silent opening beat, and an Avatar clip
per spoken part. Twelve parts with seven filmed ones lands near
<span class="credit">700</span> credits **including the re-rolls a first pass
needs** — which is the number worth quoting, because a costing that assumes
everything lands first time is not a costing.

> **Read the spend from the generation ledger, never from the balance.** On a
> shared account the balance moves for work that is not yours: 518 of an 829
> drop was one film and the rest was somebody else generating at the same time.

---

## When it breaks

### "Generate is greyed out and I do not know what is missing"

A required seat is empty. Three of the four models refuse to start without their
media, and the composer blocks rather than guessing.

> **Fix:** read the rig. Avatar wants an image and a voice track; Animate wants
> an image and a video; Replace wants a video and at least one identity. If you
> have none of those things, the model you want is P-Video.

### "The clip is a completely different length from the one I asked for"

You did not ask for one. On Avatar, Animate and Replace the length control is
not yours — the voice track or the source video sets it.

> **Fix:** change the driver. Cut a sentence out of the script, or trim the
> source video. Passing a number does nothing, and the number the composer shows
> is the fifteen-second ceiling the take is quoted at.

### "A four-second clip cost the same as a twenty-second one"

That is the flat price working as designed, and it is usually good news.

> **Fix:** nothing to fix — but stop splitting content across takes to keep them
> short. Write to the ceiling: one twenty-second take is a third of the price of
> three short ones, and it has no seams.

### "The face changed halfway through"

Either the piece is chained, or the sample is bad. Both happen and they look
alike.

> **Fix:** if each clip is seeded from the previous clip's last frame, that is
> generational drift and it will climb monotonically — rebuild each part from
> its own still instead. If the parts are already built from stills, re-roll the
> one that drifted, and measure the delivered clips rather than the images.
> Colour is the quickest read: a part that ends more than about 15% below the
> face chroma it opened on is drifting whatever it looks like at a glance, and
> one that opens below its own seed still was handed a used frame.

### "I asked for a gesture and it happened once, at the wrong time"

The seed pose is a beat the model returns to, not a pose it holds.

> **Fix:** treat a gesture as an opening beat and nothing more. If a hand has to
> be somewhere at a chosen second, that is an edit, not a prompt.

### "The mouth moves but the clip is silent, or laughs over nothing"

A visible performance over a track that does not carry it reads as dubbed, and
so does the reverse. It is the single artefact that tells a viewer the face was
generated.

> **Fix:** match them. If the beat is silent, the mouth stays closed — and check
> it rather than trusting it: tracing mouth-region brightness across every frame
> settles it in a way looking at six frames does not, because teeth are bright.

### "My video came back 1088 pixels wide instead of 1080"

Encoder padding. Vertical output routinely arrives eight pixels wide of the
composition it was authored against, and a 16:9 request came back 1280×704.

> **Fix:** crop four pixels each side. Do not scale — scaling moves every
> measured coordinate in the edit that follows.

### "I wrote the expression I wanted and got a different one"

On the performance models the prompt does not reach the face.

> **Fix:** put the expression in the still you seed with. If the face has to
> perform against near-silence, the job belongs on a model that responds to
> meaning rather than to a waveform, and that costs several times as much.

### "The transfer is smeared or the identity slipped"

Almost always the source video rather than the model.

> **Fix:** use a source with strong, well-lit, unambiguous motion and one
> subject in it, and a front-facing reference. A busy handheld clip with
> occlusion is the case these models are worst at, and no prompt recovers it.

---

## Worked examples

Two finished takes, each with the brief that made it. Both are a **single
generation**, and neither is several takes cut together. Only the first has a
player here — the second is a shot out of an episode that is not published, so
what it carries instead is the measurements taken off the delivered file.

Animate and Replace have no worked example at all yet. We have run Animate in
production and that take is not published either; a made-up example would be
worth less than the gap.

### A host who does not exist

_P-Video Avatar: one portrait, one voice take, no prompt._

<figure class="guide-example guide-example--wide">
<div class="guide-example-stage"><video src="/showcase/ladders/podcast-host.mp4" poster="/showcase/ladders/podcast-host.webp" width="960" height="540" controls controlslist="nodownload noplaybackrate noremoteplayback" disablepictureinpicture playsinline preload="none" aria-label="A woman at a microphone in a walnut-panelled studio talks to camera for twenty seconds about describing a picture and getting it back"></video></div>
<figcaption class="guide-example-side">
<dl class="guide-example-spec">
<div><dt>Length</dt><dd>20.16s, one take</dd></div>
<div><dt>Frame</dt><dd>16:9, 720p</dd></div>
<div><dt>Supplied</dt><dd>One portrait, one voice take</dd></div>
<div><dt>Cost</dt><dd><span class="credit">11</span> + <span class="credit">11</span> + <span class="credit">3.5</span> + <span class="credit">31.5</span></dd></div>
</dl>
<img class="guide-example-input" src="/showcase/guides/the-host-portrait.webp" width="760" height="428" loading="lazy" decoding="async" alt="The supplied portrait: the same woman at the microphone in the walnut-panelled studio, drawn as a still" />
<p class="guide-example-input-note">The portrait, made first, from an identity image made before it.</p>
</figcaption>
</figure>

A recurring host reads twenty seconds of script to camera. Nothing was
photographed and nobody was recorded: the portrait is a generated image, the
voice is a generated take, and the clip is those two files handed back to the
model.

**The prompt is empty.** Avatar takes an image, an audio track and an optional
visual prompt, and this run passed none of the third — the generation's `prompt`
field is the empty string. The visual direction it would have carried is already
settled by the portrait and the performance is driven by the audio, so **the
only text a person wrote for what you see is the text she says**.

The clip is **20.16 seconds** and the wav is 20.16 seconds. Nothing else set the
length. Six sentences with two marked pauses is twenty seconds, and the only way
to make a shorter clip is to write a shorter one.

And it cost <span class="credit">31.5</span> — the same as a four-second take,
because that is what a performance model quotes. The two still images that
decide who she is came to <span class="credit">22</span> between them, which is
most of what the clip cost: on this model the pictures are the expensive part.

<details class="guide-fold"><summary>Both briefs, in full</summary><p class="guide-fold-note">The clip has no brief: its prompt is the empty string. What follows is the voice take that drives it, and the style prompt beside it — the bracketed directions are read as performance notes and are not spoken.</p><pre><code>[warm, conversational, speaking at a natural pace]&#10;I type what I'm picturing, and it comes back real — images, clips, music. [short pause] Want a whole video? The assistant storyboards it, same character in every scene. [short pause] Voiced, scored, done — that's just picturing.</code></pre><pre><code>Speak as a warm, reassuring documentary narrator. Unhurried pace, gentle&#10;emphasis, close and friendly.</code></pre></details>

### Three seconds of nothing happening

_P-Video: one Start frame, no dialogue, no source footage._

A woman watches something on a monitor and says nothing. It is the opening beat
of a longer piece, and it is the shot that made the case for this model: no line
meant no audio, which left Avatar with no length and no performance to build
from, and there was no footage for Animate or Replace to recast. **P-Video was
the only one of the four that could take it.**

It came back **1088×1920, 24fps, 97 frames, 4.042 seconds**, for
<span class="credit">13.5</span> at 1080p.

Three things the take settles, each measured on the delivered file rather than
watched:

**The mouth never opens.** The first attempt had her laugh, which plays against
a silent track and reads as dubbed. Mouth-region 99th-percentile brightness
across all 97 frames — teeth are bright, so an opening mouth spikes it — reads
`198 199 200 201 202 202 203 230 230 230 229 228 232` on the laughing take and
`198 200 200 199 198 198 199 201 200 199 199 198 197` on the one we kept. The
first shows teeth from frame 56 onward. The second never leaves its baseline.

**The expression builds early and holds**, rather than arriving as a late
payoff, which is what makes the shot safe to trim from either end. The version
that laughed had to be cut 0.8 → 3.8 to protect a moment at 2.0s; this one has
no moment to protect, so 0.0 → 3.0 works and the trim comes off the tail.

**The camera stayed locked, and that was checked rather than assumed.** Mean
absolute difference against frame 0 runs 1.19–1.43 on a bare patch of wall and
1.44–1.71 on a shelf full of hard edges — both encoder noise, and separated by
about a third of a grey level. A push-in would have separated those two rows by
an order of magnitude. Nothing in the set moved but her.

<blockquote class="guide-quote--caution">
<p><strong>A verification step that can be wrong in the alarming direction is worse than none, because it is believed.</strong> An early check compared a 1152×2048 still against a 1088×1920 frame without rescaling, walked two different row strides against each other, and reported that the model had repainted the opening frame — on three clips that were correct. Rescale before comparing frames.</p>
</blockquote>

---

## About justpictur.ing

justpictur.ing is an AI picture studio: prompt-driven images, clips, voice-overs and music, and storyboard-backed films with consistent characters across every scene.

Agents connect over MCP and can generate directly:

```
claude mcp add justpicturing https://api.justpictur.ing/mcp --transport http
```

The app is at https://app.justpictur.ing and requires an account.

Every guide on this site has a Markdown twin at its own URL plus `.md`. Every guide: https://justpictur.ing/guides/
