ByteDance · Clips · premium
OmniHuman 1.5
ByteDance's performance model: a portrait and an audio track, at 1080p with real body motion.
What it costs
| 5s | 67 credits |
What it takes
| Aspect ratios | 16:9 4:3 1:1 3:4 9:16 21:9 |
| Resolutions | 1080p |
| Duration | 1–15 seconds |
| Reference images | 1 |
| Reference audio | 1 |
| Audio | Always generated |
| Required inputs | image, audio |
When to reach for it
The high-fidelity version of the talking-clip job. Like P-Video Avatar it needs an image and an audio track and animates the person to speak, but it renders at 1080p only, and it moves considerably more than the mouth. Gesture, posture and head motion all respond to the audio, which is the difference between a clip that reads as a performance and one that reads as a photo with a moving jaw.
It costs several times what Avatar does for the same length, and that is the whole decision. For a line of dialogue inside a longer sequence, or anything where the person is the subject rather than a talking element in a corner, the difference is visible enough to justify it. For volume work, twenty short lines or a placeholder or a face in a montage, it is not.
Duration follows the audio up to fifteen seconds. One reference image, one reference audio, and audio is always generated. Six aspect ratios, so vertical and square both work.
Compared with
- P-Video Avatar vs OmniHuman 1.5
Avatar for volume and for faces in the background. OmniHuman when the person is the subject.
OmniHuman 1.5 is one pick in the composer.
Every model in this catalog runs from the same prompt bar, priced in the same credits. New accounts start with some to spend.
Start picturing