NewSeedance 2.5just landed. A full 30 seconds, with sound, in one take.30 seconds, with sound, one take.Try it immediatelyTry it
Guides

34 min readMarkdown

P-Video: the complete guide to Pruna's four models

Four models, one family: a cheap generator and three performance rigs. What each one takes, what a take really costs, and the failures worth knowing first.

Pruna makes four video models, and they do not compete with each other. One makes a clip out of a prompt. The other three animate something you already have: a portrait and a voice, a still and a performance, a shot and a new face in it.

They have become the models we reach for first, and the reason is arithmetic. Twenty seconds of someone talking to camera costs 31.5 here and 200 on the model above it. Five seconds of 720p costs 8.5, against 37.5 for the next cheapest model that will render the same shot. What comes back is good enough that the saving stops being a compromise and starts being the default.

Four models is not four things to learn. It is one question — what is already in your hands? — asked four ways.

NOTE: Everything here is true of all four unless it names one.

Overview

The short version

Four decisions do most of the work. If you read nothing else, read these. Each one is written out in full further down.

  1. Pick the model from what you already have, not from what you want

    The four split cleanly by input and there is no overlap between them. Three of them refuse to start without media you supply, so the model is chosen for you the moment you know whether you are holding a voice, a video, or nothing.

    Nothing but words → P-Video. A portrait and a voice → Avatar. A still and a performance to copy → Animate. A shot whose person should be somebody else → Replace.

  2. A performance take is a flat price, whatever comes back

    Avatar, Animate and Replace read their length off the media you attached, so the take is quoted at the fifteen-second ceiling and never above it. A 4.5-second clip and a 17.9-second one came to the same number — so a short line wastes the budget a long one would have used for free.

    Write to the ceiling. Two sentences and six sentences cost the same, so put the whole thought in one take rather than splitting it across two.

  3. What you attach outranks what you write

    On the three performance models the driver carries the timing, the motion and the performance, and the prompt only decorates what is left. A property the media sets cannot be argued out of the model in prose — asked for "curious, then a chuckle, then a complicated smile", the model returned one continuous pleasant smile, because it drives the mouth from the waveform.

    Instead of

    Writing the expression you want into the prompt

    Do

    Put it in the still you seed with, or in the take you drive with

  4. P-Video has no reference slot — it has frames

    You cannot hand it a character and ask for that character. What you can hand it is the exact picture the clip opens on, the exact picture it arrives at, or both. A supplied frame also overrules your The shape of the frame: its width against its height. 16:9 is the wide shape a television is; 9:16 is a phone held upright. setting, because the ratio comes off the image.

    Settle the subject in one image first, then hand that image over as the Start frame. Consistency comes from the picture here, never from the prompt.

The family, in one table

Model Needs Also takes Length Frame shape Sound
P-Video a prompt Start frame, End frame, one audio track 1–10s, or the audio’s seven, unless a frame sets it generated, switchable
P-Video Avatar a portrait and a voice track an optional prompt the voice’s, exactly the portrait’s the voice you gave it
P-Video Animate an image and a source video an optional prompt the source’s the source’s the source’s, always on
P-Video Replace a source video and 1–3 identities an optional prompt the source’s the source’s the source’s, switchable

All four render at 720p or 1080p, and all four take a prompt of up to 4,000 characters — twice what an image model gets, and more than any of them needs. On three of the four the prompt is optional, and one of the worked examples below ran with it empty.

What it costs

P-Video is priced by the second, and it is the cheapest way in the catalog to turn a prompt into a moving picture — $0.02 a second at 720p, under half the rate of anything else that renders at that size.

Length 720p 1080p What it’s for
1s 2 3.5 A punctuating cut. One second is a real option here
3s 5 10 A reaction, a beat, a shot with one thing happening in it
5s 8.5 17 The default. A complete idea
10s 17 33.5 The ceiling, and long enough to hold two beats

The other three are a flat price. The length belongs to the media you attached, so no duration is ever sent and the take is quoted at the ceiling — fifteen seconds’ worth — with settlement never going above it. Every take we have run has come to exactly this number, whether the clip was four seconds or twenty:

Model 720p 1080p
P-Video Avatar 31.5 56.5
P-Video Animate 37.5 75
P-Video Replace 37.5 75

Fifteen seconds is where the price stops, not where the clip stops. The composer shows the ceiling because it is what gets held, and the model never receives a duration at all — a 17.9-second line and a 20.16-second one both came back whole, and both cost 31.5.

So do not split a part to fit a number that is not a limit. Split it at a real seam, or not at all.

Why they became the default

The same three jobs, priced against what else in the catalog can do them:

The job Pruna What else does it
Five seconds of 720p clip P-Video 8.5 Seedance 2.0 Mini 37.5 · Seedance 2.5 at 480p 43 · Runway Gen-4.5 50
Twenty seconds of someone talking P-Video Avatar 31.5 OmniHuman 1.5 200 · three 8s Veo 3.1 Fast takes joined 300 · Seedance 2.5 at 720p 385.5
Recasting a performance onto a character P-Video Animate 37.5 Kling Motion Control 87.5

These are not the best models in the catalog, and the guide does not claim they are. What they are is the models where the extra spend stops buying anything a viewer notices — and that is a different, more useful property.

Reach past them when the face is the performance rather than the thing the words come out of, when the camera has to do something cinematic, or when a generated shot needs to run longer than ten seconds in one piece.


The rules

Seven things that are true whatever you are making. Each one cost us money to learn, and they are ordered by how much they will save you.

Choose from your inputs, not from your ambition

Three of the four refuse to start. What they refuse over is the model choice, made for you.

A generation here is blocked until its required seats are full, and that is the whole taxonomy:

  • Avatar wants an image and an audio track.
  • Animate wants an image and a video.
  • Replace wants a video and between one and three identities.
  • P-Video wants nothing at all.

We hit this on a reaction shot: three seconds of a woman watching something, saying nothing. No line meant no audio, which meant no length and no performance for a lip-sync model to build from; there was no footage to recast either. P-Video was the only one of the four that could take the shot, and it took it for 13.5 at 1080p.

Reading the requirement backwards is the fast route to the right model. If you are reaching for Avatar and have no voice take yet, the job is not “generate a clip”, it is “cut the voice first”.

A performance take is a flat price, so write to the ceiling

The clip is as long as what drives it, and the bill is as long as fifteen seconds either way.

None of the three performance models takes a length. Every request is the auto sentinel, no number ever reaches the provider, and the take is quoted at the fifteen-second ceiling regardless of what the media will turn out to be worth. Measured across our own runs:

  • A 4.48-second avatar clip and a 17.9-second one both came to 31.5 at 720p.
  • A twenty-second one, in the worked example below, came to the same 31.5.

That inverts the usual instinct. On a per-second model you write short to save money; here writing short wastes the seconds you already paid for. Six sentences with two marked pauses is twenty seconds and one take. The same content split across two takes is twice the price and a cut you have to hide.

The corollary is the budgeting rule: a lip-synced piece is budgeted from the script, not from the clip. The only way to make a shorter take is to write a shorter line.

The driver is the clock, so cut it first

The audio decides the length, which means the audio has to exist before the picture does.

On Avatar this is exact rather than approximate. A 967,724-byte 24kHz mono wav is 20.16 seconds and the clip came back 20.16 seconds. Across four more takes each clip landed within 50 milliseconds of the voice-over that drove it, and the delivered soundtrack correlated with the source mp3 at 0.9997 — the model re-emits the track you gave it rather than resynthesising anything.

So the running order is voice, then picture, in that order, every time. Do it the other way round and you are trimming a face to fit a line.

The same rule holds one step sideways on Animate and Replace, where the source video is the clock: its length, its timing, its camera and its frame rate all survive into the output, and the only way to change any of them is to change the source.

What the picture sets, the prompt cannot argue out

A property the input carries is not up for discussion downstream.

This is the single most expensive thing to learn about the family, and it shows up in three different places.

On expression. Given “curious, then a chuckle, then a complicated smile”, Avatar returns one continuous pleasant smile. It drives the mouth from the waveform and ignores the prompt’s semantics. When the face is the performance and the audio is nearly silent, that is what OmniHuman 1.5 costs 200 for.

On gesture. A hand in the seed still is a beat the model returns to, not a pose it holds. On a ten-second take a raised index finger was up at 0.2s and 1.5s, gone by 3s, back at 7s, gone by 9s. On takes under about six seconds the hands play once at the top and settle out of frame. So a gesture still buys you a reliable opening beat and nothing you can place — never write a shot that needs a hand present at a chosen moment.

On framing. The output takes its shape from the portrait or from the source video, not from the ratio you selected. A 16:9 request came back 1280×704, and a vertical take came back 1088×1920 — eight pixels of encoder padding on a 1080-wide composition. The setting shapes the composer’s preview; the file is the authority. Crop, do not scale, or every measured coordinate in your edit moves.

Build from stills, not from a chain

Each hop re-renders the last one, and the drift is monotonic.

It is tempting to run a sequence by feeding each clip’s final frame into the next: the boundaries carry no visible cut, and it is one fewer image to make. We built a four-part sequence that way and measured face-region difference against the original anchor at each step:

29.2 → 35.2 → 36.9 → 41.2.

A clean climb, which is what generational drift looks like. Rebuilt from a fresh still per part, the same measure reads 35.6 · 40.2 · 43.6 · 32.7 — unordered, because what it now tracks is pose rather than accumulation. Chaining also freezes one pose for the whole piece: she never gestures, never sits back, never does anything but talk from where the previous clip left her.

Chain only where the performance genuinely continues. One boundary in that sequence earned it — she is still smiling at the screen when the next line starts, and she lifts her eyes on it. Frame 0 of the second clip differed from the first clip’s final frame by 3.78/255, against a bare-wall patch that could not have changed at 3.99. Same magnitude: what separates them is re-encode noise and nothing else.

Everywhere else, a fresh still is a hard cut and the only way a shot gets its own pose. Cheap, too — one image is a fraction of the clip it seeds.

The drift has a second signature, and this one you can see. Face difference is a number; colour is a picture. Mean chroma across a fixed crop of the face falls inside every take that runs long enough, and a continuation opens on the value the previous take closed at — so the fall carries across the boundary instead of resetting at it. Two takes from one session, both 9:16, one identity, four minutes apart:

Two rows of six face crops of the same woman. The top row is warm-skinned and freckled; the bottom row is grey, waxy and hollow-cheeked.

Top: a take seeded from its own still, frames 80–110. Bottom: a take that opened on the previous clip's final frame, frames 84–114. Same wardrobe, same wall, same light, four minutes apart.

The top take loses 8.6% of its face colour over 4.7 seconds and is still recognisably her at the end of it. The bottom take is one hop further on. It opens 28% below where the top one opens, closes 44% below, and spends 22.6% inside its own 4.9 seconds — the steepest of the nine takes we measured that night, and the only one to finish below 60% of where the top take opened. What that number looks like is the picture: pallor, waxy skin with the pore detail gone, hollow cheeks, sunken sockets, an older face.

The handoff is literal. Frame 0 of the failing take and the final frame of the take before it are not similar, they are the same picture — 36.9 dB PSNR, against 16.7 dB for an unrelated pair from the same session.

Four face crops: a take's first and last frame, then the next take's first and last frame. The middle two are identical; the fourth is grey and drawn.

The parent take's first and last frame, then the child's. The middle two are one picture, handed over. The fourth is 4.9 seconds after it.

The still sets the budget, and the take spends it. A first-hop take opens within 4% of the colour of the still it was seeded from — in all four pairs we could match, the model preserved what it was given and then bled it. So a chain does not go wrong at the end. It goes wrong at the picture you hand it, and every hop after that spends what is left:

Take Opened on Length Opens Closes
1 its own still 3.0s 105% 99%
2 its own still 3.0s 101% 103%
3 its own still 2.9s 95% 96%
4 its own still 5.5s 103% 87%
5 its own still 6.7s 92% 75%
6 its own still 4.7s 100% 91%
7 take 6’s final frame 10.6s 94% 86%
8 its own still 10.6s 79% 71%
9 take 8’s final frame 4.9s 72% 56%

Nine takes, one session, one identity, one room. Each figure is mean face-region chroma over the take’s first and last tenth, against take 6’s opening. Two things spend the budget and they are not the same thing: length — three seconds costs almost nothing, ten costs about a tenth — and hops. Take 7 chains and survives it, because it opened at 94%. Take 9 chains from a take that had already opened low, and lands somewhere no prompt and no re-roll recovers.

Which makes it cheap to catch. One pass of ffmpeg’s signalstats over a crop of the face reports chroma per frame, and two thresholds separate the good takes from the corpse here without anything cleverer: a take that closes more than about 15% below where it opened is drifting, and a take that opens below the still it was seeded from has inherited a used frame.

It will animate a photographed face

The refusal you hit elsewhere is not here, and for a lot of work that is the whole reason to be here.

Hand Seedance 2.5 a photorealistic close-up of a person and it declines — seven refusals in seven attempts in our tests, through every door. Pruna’s models take the same picture without comment.

That is the axis on which the choice actually turns for User-generated content: an ad shot to look like a real customer filmed it on their phone, rather than like an ad. work: if your piece needs a believable person talking to camera from a picture you already have, this family will do it. So will Veo 3.1, which is the stronger speech model — but Veo stops at eight seconds and charges 100 to get there on its Fast tier, where twenty seconds here is 31.5.

It is not a licence. A face you have no right to use is still a face you have no right to use. Every person in every example on this page was generated, and nothing was photographed.

Test at 720p, and let the price do the testing

1080p costs double on three of the four, and a little under double on Avatar. Nothing else about a take changes.

There is no rehearsal model here and no draft tier — the provider has one, and the catalog deliberately does not wire it up, because the full-quality run is already cheaper than most models’ rehearsals. Five seconds of P-Video at 720p is 8.5; the same brief at 1080p is 17. Dropping a tier changes sharpness and changes nothing a prompt controls.

So the working method is: run it at 720p and read the result honestly. It answers everything a brief can get wrong — whether the framing is what you pictured, whether the motion holds, whether your inputs are accepted, whether the character survives. Then re-run the same words at 1080p for the version you keep.

One provider note that cuts the other way: vertical work is reported to hold up better at 1080p. If a 9:16 take looks soft in a way 720p usually does not explain, that is the first thing to try.


The four models

Each one is a different job. They are written here in the order you meet them: the one that needs nothing, then the three that need something.

P-Video — the one that needs nothing

What it is for

  • Best for: Footage with nobody speaking in it, like scenery, hands, a street or a detail, cut in around the main shots to cover a voice-over or to let a scene breathe., close-ups, product animation, reaction shots with no line in them, and the connective tissue of a longer piece.
  • Costs you: reference images. There is no slot for one, so a subject cannot be carried in.
  • Watch for: the aspect ratio comes from a supplied frame, not from your setting.

What it takes

Seven aspect ratios, one to ten seconds, 720p or 1080p, and three optional inputs that change what the model is doing:

Start frame     the clip opens on this exact picture
End frame       the clip arrives at this exact picture
Audio track     the clip is cut to this, and takes its length from it

Attach an audio track and the duration control goes away: the sound decides, exactly as it does on the performance models. That is the cheapest music-video route in the catalog, and it is unusual to find an audio input at this price at all.

Sound is generated by default and costs nothing extra. Unlike Veo 3.1, where the audio switch doubles the rate, speech and effects here come out of the same per-second price whether you write for them or not. The provider is candid that sound effects are the weaker half — for anything where the sound is the product, make the track in the Audios room and hand it back as the audio input.

How to write it

The provider’s own formula is a good starting shape, and only the first three parts are required:

[Subject] [Action] [Scene] [Camera] [Lighting] [Style] [Audio]

Two habits carry over from the rest of this shelf and matter more than the formula does. Name what is visible in the frame rather than the camera term for it — a description of what should be on screen gets followed where an angle gets interpreted. And describe anyone who has to stay the same once, in detail, then point back to that; describing them twice invites a redraw.

The one thing the prompt cannot do is hold a subject across takes. Nothing is carrying identity, so if the same person has to appear in two clips, settle them in an image first and hand that image in as the Start frame.

The known limits

The model’s own README is unusually direct about where it stops, and our runs agree: no extreme cinematic camera motion, no native 4K, weak sound effects, and speaker attribution that drifts above two voices. It is very good at close-up subjects and foreground objects, and it is not a model to plan a multi-scene story around.

Worked example: Three seconds of nothing happening — a reaction shot no other model in the family could take.

P-Video Avatar — a portrait and a voice

What it is for

  • Best for: anything scripted to camera — a host, an explainer, a UGC ad, a localised version of a piece you already made.
  • Costs you: the choice of length, and any control over the performance beyond what the still and the track already carry.
  • Watch for: it drives the mouth from the waveform, so a silent moment is not a performance.

What it takes

Two required seats — Avatar image and Voice audio — and an optional prompt. That is the whole recipe. The audio’s voice and timing animate your image, and the clip is as long as the track — to the frame, and past fifteen seconds without paying for it.

The provider will also synthesise speech itself, from a script and one of thirty named voices. We deliberately do not wire that up. The Audios room already writes voice-overs across five TTS models, so a take from there is the avatar’s voice, and there is one voice picker in the product rather than two. It also makes the lineage legible: the voice exists as its own generation, with its own prompt and its own cost, and the script can be rewritten and re-cut without touching the host at all.

How to write it

This is the model where less prompt is better, and often none at all is best.

The two required inputs already specify almost everything: who is talking and what they are saying. The prompt is visual direction on top of that, and the portrait has usually settled the visual direction already.

  • Use a clear, front-facing portrait. Heavy angles, occlusion and low resolution all cost you identity.
  • Keep the audio clean. Clear speech with little background noise gives tighter sync.
  • Put the performance in the still, not in the prompt — see the rule above.

The known limit

It will not play a scripted expression arc. The mouth follows the waveform and the semantic content of the prompt does not reach the face. It is right when the line carries the performance; when the face carries it and the audio is nearly silent, this is not the model.

One more, and it is the kind you only catch by looking: the face drifts on some samples, and the still is not at fault. Two clips in one run came back narrower and harder with patchier skin while all six source stills measured identical. A re-roll fixed it. Check the delivered clips, not the images.

Worked example: A host who does not exist: twenty seconds, two inputs, and an empty prompt.

P-Video Animate — a performance, recast

What it is for

  • Best for: putting a character you designed into a movement you already have — a walk, a dance, a piece of found choreography, a performance you filmed once and want to reuse.
  • Costs you: any control the source video already exercises. The pace, the camera, the timing and the sound all arrive with it.
  • Watch for: the image is identity, not frame zero. It is not a Start frame and it will not be treated as one.

What it takes

Image to animate plus Driving performance video, and an optional prompt. The subject of your image performs the motion of the source, and the source’s audio is muxed onto the output — always. There is no switch, which is why the model reads as sound-required in the composer.

That is the difference between this and generating the same shot from a prompt. A motion transfer proves you can move a performance onto somebody else; a generated version proves the model can invent a walk. Both are true statements about the product, and only one of them needs footage.

How to write it

The prompt is direction on top of a performance. It is not the performance.

This is why an Animate brief is four sentences where the equivalent prompt-only brief runs four paragraphs: everything a prompt would otherwise have to choreograph — the pace, the arm that stays in the pocket, the coat hem, the gaze arriving on the last two steps — is already in the source and does not need writing down.

What the source cannot supply is what to preserve during the transfer, and that is what the prompt is for: performance emphasis, and what must remain unchanged.

  • Strong, well-lit motion transfers. Subtle or occluded movement does not, and a busy handheld source is the usual reason a transfer fails.
  • A front-facing reference image preserves identity; heavy angles do not.
  • Frame the source the way you want the result framed, because the output keeps the source’s aspect ratio and its frame rate.

Where to start

The catalog has a published recipe on this model — Fashion Walk / Outfit Reel — which takes a styled-look image and a walking performance and gives you the two seats already labelled. It is the fastest way to see the shape of the job before you build one of your own.

P-Video Replace — the same shot, a different person

What it is for

  • Best for: keeping a shot and changing who is in it — ad variants across markets, a mascot dropped into existing footage, one strong performance reused with different talent.
  • Costs you: the scene. This model does not invent one; it preserves the one you gave it.
  • Watch for: the result is only as good as the match between your subject’s proportions and the original’s.

What it takes

Source video plus one to three Replacement identities, and an optional prompt. Motion, timing, camera and scene all stay exactly as they were; the person changes. The source’s frame rate is pinned to original, so nothing is resampled behind your back.

It is the only one of the three performance models where the sound is still a choice — Animate welds the source track on, and this one leaves the switch live.

How to write it

Three identity slots rather than one is the detail that matters, and it is worth spending them. Three genuinely different views of the same subject is what makes a replacement hold across a shot with camera movement in it; three frames of the same pose is one reference with extra steps.

Otherwise the guidance is Animate’s, for the same reasons: clear, front-facing references with minimal occlusion, and strong, well-lit motion in the source. The prompt’s job is to say how the people from the references should sit in the scene, and nothing more.

Putting them together

The four are worth more as a chain than as four models, and the chain is short:

an image model      →  a still that settles who and where
P-Video             →  the shots with nobody speaking
P-Video Avatar      →  the shots with a line, one still and one voice take each
P-Video Animate     →  a shot you already have a performance for
P-Video Replace     →  a shot you already have entirely

A twelve-part vertical short we built runs almost exactly that way: one anchor image holding the character, a fresh still per part changing only the eyeline and the hands, one P-Video clip for the silent opening beat, and an Avatar clip per spoken part. Twelve parts with seven filmed ones lands near 700 credits including the re-rolls a first pass needs — which is the number worth quoting, because a costing that assumes everything lands first time is not a costing.

Read the spend from the generation ledger, never from the balance. On a shared account the balance moves for work that is not yours: 518 of an 829 drop was one film and the rest was somebody else generating at the same time.


When it breaks

“Generate is greyed out and I do not know what is missing”

A required seat is empty. Three of the four models refuse to start without their media, and the composer blocks rather than guessing.

Fix: read the rig. Avatar wants an image and a voice track; Animate wants an image and a video; Replace wants a video and at least one identity. If you have none of those things, the model you want is P-Video.

“The clip is a completely different length from the one I asked for”

You did not ask for one. On Avatar, Animate and Replace the length control is not yours — the voice track or the source video sets it.

Fix: change the driver. Cut a sentence out of the script, or trim the source video. Passing a number does nothing, and the number the composer shows is the fifteen-second ceiling the take is quoted at.

“A four-second clip cost the same as a twenty-second one”

That is the flat price working as designed, and it is usually good news.

Fix: nothing to fix — but stop splitting content across takes to keep them short. Write to the ceiling: one twenty-second take is a third of the price of three short ones, and it has no seams.

“The face changed halfway through”

Either the piece is chained, or the sample is bad. Both happen and they look alike.

Fix: if each clip is seeded from the previous clip’s last frame, that is generational drift and it will climb monotonically — rebuild each part from its own still instead. If the parts are already built from stills, re-roll the one that drifted, and measure the delivered clips rather than the images. Colour is the quickest read: a part that ends more than about 15% below the face chroma it opened on is drifting whatever it looks like at a glance, and one that opens below its own seed still was handed a used frame.

“I asked for a gesture and it happened once, at the wrong time”

The seed pose is a beat the model returns to, not a pose it holds.

Fix: treat a gesture as an opening beat and nothing more. If a hand has to be somewhere at a chosen second, that is an edit, not a prompt.

“The mouth moves but the clip is silent, or laughs over nothing”

A visible performance over a track that does not carry it reads as dubbed, and so does the reverse. It is the single artefact that tells a viewer the face was generated.

Fix: match them. If the beat is silent, the mouth stays closed — and check it rather than trusting it: tracing mouth-region brightness across every frame settles it in a way looking at six frames does not, because teeth are bright.

“My video came back 1088 pixels wide instead of 1080”

Encoder padding. Vertical output routinely arrives eight pixels wide of the composition it was authored against, and a 16:9 request came back 1280×704.

Fix: crop four pixels each side. Do not scale — scaling moves every measured coordinate in the edit that follows.

“I wrote the expression I wanted and got a different one”

On the performance models the prompt does not reach the face.

Fix: put the expression in the still you seed with. If the face has to perform against near-silence, the job belongs on a model that responds to meaning rather than to a waveform, and that costs several times as much.

“The transfer is smeared or the identity slipped”

Almost always the source video rather than the model.

Fix: use a source with strong, well-lit, unambiguous motion and one subject in it, and a front-facing reference. A busy handheld clip with occlusion is the case these models are worst at, and no prompt recovers it.


Worked examples

Two finished takes, each with the brief that made it. Both are a single generation, and neither is several takes cut together. Only the first has a player here — the second is a shot out of an episode that is not published, so what it carries instead is the measurements taken off the delivered file.

Animate and Replace have no worked example at all yet. We have run Animate in production and that take is not published either; a made-up example would be worth less than the gap.

A host who does not exist

P-Video Avatar: one portrait, one voice take, no prompt.

Length
20.16s, one take
Frame
16:9, 720p
Supplied
One portrait, one voice take
Cost
11 + 11 + 3.5 + 31.5
The supplied portrait: the same woman at the microphone in the walnut-panelled studio, drawn as a still

The portrait, made first, from an identity image made before it.

A recurring host reads twenty seconds of script to camera. Nothing was photographed and nobody was recorded: the portrait is a generated image, the voice is a generated take, and the clip is those two files handed back to the model.

The prompt is empty. Avatar takes an image, an audio track and an optional visual prompt, and this run passed none of the third — the generation’s prompt field is the empty string. The visual direction it would have carried is already settled by the portrait and the performance is driven by the audio, so the only text a person wrote for what you see is the text she says.

The clip is 20.16 seconds and the wav is 20.16 seconds. Nothing else set the length. Six sentences with two marked pauses is twenty seconds, and the only way to make a shorter clip is to write a shorter one.

And it cost 31.5 — the same as a four-second take, because that is what a performance model quotes. The two still images that decide who she is came to 22 between them, which is most of what the clip cost: on this model the pictures are the expensive part.

Both briefs, in full

The clip has no brief: its prompt is the empty string. What follows is the voice take that drives it, and the style prompt beside it — the bracketed directions are read as performance notes and are not spoken.

[warm, conversational, speaking at a natural pace]
I type what I'm picturing, and it comes back real — images, clips, music. [short pause] Want a whole video? The assistant storyboards it, same character in every scene. [short pause] Voiced, scored, done — that's just picturing.
Speak as a warm, reassuring documentary narrator. Unhurried pace, gentle
emphasis, close and friendly.

Three seconds of nothing happening

P-Video: one Start frame, no dialogue, no source footage.

A woman watches something on a monitor and says nothing. It is the opening beat of a longer piece, and it is the shot that made the case for this model: no line meant no audio, which left Avatar with no length and no performance to build from, and there was no footage for Animate or Replace to recast. P-Video was the only one of the four that could take it.

It came back 1088×1920, 24fps, 97 frames, 4.042 seconds, for 13.5 at 1080p.

Three things the take settles, each measured on the delivered file rather than watched:

The mouth never opens. The first attempt had her laugh, which plays against a silent track and reads as dubbed. Mouth-region 99th-percentile brightness across all 97 frames — teeth are bright, so an opening mouth spikes it — reads 198 199 200 201 202 202 203 230 230 230 229 228 232 on the laughing take and 198 200 200 199 198 198 199 201 200 199 199 198 197 on the one we kept. The first shows teeth from frame 56 onward. The second never leaves its baseline.

The expression builds early and holds, rather than arriving as a late payoff, which is what makes the shot safe to trim from either end. The version that laughed had to be cut 0.8 → 3.8 to protect a moment at 2.0s; this one has no moment to protect, so 0.0 → 3.0 works and the trim comes off the tail.

The camera stayed locked, and that was checked rather than assumed. Mean absolute difference against frame 0 runs 1.19–1.43 on a bare patch of wall and 1.44–1.71 on a shelf full of hard edges — both encoder noise, and separated by about a third of a grey level. A push-in would have separated those two rows by an order of magnitude. Nothing in the set moved but her.

A verification step that can be wrong in the alarming direction is worse than none, because it is believed. An early check compared a 1152×2048 still against a 1088×1920 frame without rescaling, walked two different row strides against each other, and reported that the model had repainted the opening frame — on three clips that were correct. Rescale before comparing frames.

Everything here runs in the studio.

Every model named on this page is one you can pick from the composer. New accounts start with credits to spend on exactly this.

Start picturing