Pruna AI · Clips · standard
P-Video Avatar
A photo plus an audio track becomes a talking clip. The prompt is optional.
What it costs
| 5s · 720p | 10.5 credits |
| 5s · 1080p | 19 credits |
What it takes
| Aspect ratios | 16:9 4:3 1:1 3:4 9:16 21:9 |
| Resolutions | 720p 1080p |
| Duration | 1–15 seconds |
| Reference images | 1 |
| Reference audio | 1 |
| Audio | Always generated |
| Required inputs | image, audio |
When to reach for it
A performance model rather than a generator: it requires an image and an audio track, and it animates the person in the image to speak the audio. The prompt is optional, because the two required inputs already specify almost everything: who is talking and what they are saying.
That makes it the cheapest route to a talking-head clip in the catalog by a wide margin. Ten and a half credits for a five-second 720p take, nineteen at 1080p. Pair it with one of the voice models and a single portrait, and a scripted piece to camera costs very little per line.
Duration is set by the audio rather than chosen: it runs as long as the track, up to fifteen seconds. That is the right behaviour for lip sync and it means the length control in the composer is not yours to set. For higher fidelity on the same job, OmniHuman 1.5 does it at 1080p with better motion and costs several times as much; for syncing to existing *footage* rather than a still, Kling Lip Sync is the one.
Made with it

The same host, every episode
A face you keep. A voice you cast.
Compared with
- P-Video Avatar vs OmniHuman 1.5
Avatar for volume and for faces in the background. OmniHuman when the person is the subject.
Guides that cover it
- P-Video: the complete guide to Pruna's four models
Four models, one family: a cheap generator and three performance rigs. What each one takes, what a take really costs, and the failures worth knowing first.
P-Video Avatar is one pick in the composer.
Every model in this catalog runs from the same prompt bar, priced in the same credits. New accounts start with some to spend.
Start picturing