NewSeedance 2.5just landed. A full 30 seconds, with sound, in one take.30 seconds, with sound, one take.Try it immediatelyTry it

ByteDance · Clips · premium

OmniHuman 1.5

ByteDance's performance model: a portrait and an audio track, at 1080p with real body motion.

What it costs

5s67 credits

What it takes

Aspect ratios16:9 4:3 1:1 3:4 9:16 21:9
Resolutions1080p
Duration1–15 seconds
Reference images1
Reference audio1
AudioAlways generated
Required inputsimage, audio

When to reach for it

The high-fidelity version of the talking-clip job. Like P-Video Avatar it needs an image and an audio track and animates the person to speak, but it renders at 1080p only, and it moves considerably more than the mouth. Gesture, posture and head motion all respond to the audio, which is the difference between a clip that reads as a performance and one that reads as a photo with a moving jaw.

It costs several times what Avatar does for the same length, and that is the whole decision. For a line of dialogue inside a longer sequence, or anything where the person is the subject rather than a talking element in a corner, the difference is visible enough to justify it. For volume work, twenty short lines or a placeholder or a face in a montage, it is not.

Duration follows the audio up to fifteen seconds. One reference image, one reference audio, and audio is always generated. Six aspect ratios, so vertical and square both work.

Compared with

OmniHuman 1.5 is one pick in the composer.

Every model in this catalog runs from the same prompt bar, priced in the same credits. New accounts start with some to spend.

Start picturing