Models
42 models, and what each one is for.
What each one takes, what it makes, and what a run costs in credits, read from the same table the composer charges against. Pick one and its own page carries the full price list.
Images
Seedream 5 LiteByteDanceBest value
The default image model: cheap, fast, takes seven references, and gives away its top resolution.
A fast, inexpensive all-rounder that still takes seven references: the everyday default for product and lifestyle shots.
- Aspect ratios
- 1:1 4:3 3:4 16:9 9:16 3:2 2:3 21:9
- Reference images
- 7
- Cheapest run
- from 3 credits
Nano Banana 2Google
Google's image model, and the one to reach for when the aspect ratio is unusual.
Google's flagship image model: crisp legible text, faithful prompt-following, and clean edits from up to seven references.
- Aspect ratios
- 1:1 1:4 1:8 2:3 3:2 3:4 4:1 4:3 4:5 5:4 8:1 9:16 16:9 21:9
- Reference images
- 7
- Cheapest run
- from 6 credits
GPT Image 2OpenAI
OpenAI's image model. The best instruction-follower here, with a quality dropdown that swings the price elevenfold.
OpenAI's image model: strong instruction-following and precise, mask-free edits when you need exactly what you described.
- Aspect ratios
- 1:1 3:2 2:3 4:3 3:4 16:9 9:16
- Reference images
- 7
- Cheapest run
- from 1 credit
Nano Banana 2 LiteGoogle
Nano Banana's fourteen aspect ratios at half the price, capped at 1K.
The quickest, cheapest Nano Banana: the same look, trimmed for speed while you iterate.
- Aspect ratios
- 1:1 1:4 1:8 2:3 3:2 3:4 4:1 4:3 4:5 5:4 8:1 9:16 16:9 21:9
- Reference images
- 7
- Cheapest run
- from 3 credits
Seedream 5 ProByteDance
Seedream with more capability per pass, priced by size rather than given away.
Seedream's top tier: richer detail and stronger prompt adherence for hero images and final renders.
- Aspect ratios
- 1:1 4:3 3:4 16:9 9:16 3:2 2:3 21:9
- Reference images
- 7
- Cheapest run
- from 4 credits
FLUX SchnellBlack Forest LabsCheapest
Half a credit an image. The volume tier, for when you need twenty of something.
Black Forest Labs' speed build: pennies per image for quick text-to-image drafts and thumbnails.
- Aspect ratios
- 1:1 2:3 3:2 3:4 4:3 4:5 5:4 9:16 16:9 21:9
- Reference images
- Not accepted
- Cheapest run
- from 0.5 credits
FLUX.2 FlexBlack Forest Labs
FLUX.2 with more control per pass, priced a step above Pro.
FLUX.2 with tunable quality-vs-cost: dial resolution up for finals or down to sketch cheaply.
- Aspect ratios
- 1:1 2:3 3:2 3:4 4:3 4:5 5:4 9:16 16:9
- Reference images
- 7
- Cheapest run
- from 2.5 credits
FLUX.2 ProBlack Forest Labs
Black Forest Labs' workhorse, and one of the cheapest reference-capable models here.
The production FLUX.2: dependable, detailed renders with reference support at a flat rate.
- Aspect ratios
- 1:1 2:3 3:2 3:4 4:3 4:5 5:4 9:16 16:9
- Reference images
- 7
- Cheapest run
- from 2 credits
FLUX.2 MaxBlack Forest Labs
The top of the FLUX.2 ladder, for when the render itself is the deliverable.
The most capable FLUX.2: maximum fidelity and prompt control for demanding compositions.
- Aspect ratios
- 1:1 2:3 3:2 3:4 4:3 4:5 5:4 9:16 16:9
- Reference images
- 7
- Cheapest run
- from 5 credits
RRecraft V4.1 SVGRecraft
The only way to get an editable vector out of this catalog. Logos, icons, flat illustration.
Generates real, editable vector art: logos, icons, and illustrations that scale to any size.
- Aspect ratios
- 1:1 2:3 3:2 3:4 4:3 4:5 5:4 9:16 16:9
- Reference images
- Not accepted
- Output
- SVG (editable vector)
- Cheapest run
- from 3.5 credits
RRecraft V4.1 Pro SVGRecraft
Recraft at 2048×2048: more path detail, at six times the price of the standard tier.
Recraft's pro vector tier: cleaner paths and finer detail for brand-ready SVG assets.
- Aspect ratios
- 1:1 2:3 3:2 3:4 4:3 4:5 5:4 9:16 16:9
- Reference images
- Not accepted
- Output
- SVG (editable vector)
- Cheapest run
- from 21 credits
P-ImagePruna AICheapest
Half a credit, JPEG out. The other volume tier.
Pruna's efficiency-tuned image model: quick, budget-friendly text-to-image.
- Aspect ratios
- 1:1 4:3 3:4 16:9 9:16 3:2 2:3
- Reference images
- Not accepted
- Cheapest run
- from 0.5 credits
Grok Imagine ImagexAI
xAI's image model. Two credits, and one picture in for an edit rather than a reference set.
xAI's image model: inexpensive prompt-to-image, and one picture in for an edit rather than a whole reference set.
- Aspect ratios
- 1:1 16:9 9:16 4:3 3:4 3:2 2:3
- Reference images
- 1
- Cheapest run
- from 2 credits
Grok Imagine Image QualityxAI
The sharper Grok: 2K output, better text rendering, and a resolution that is the whole decision.
The sharper Grok Imagine: finer detail and better text rendering, with 2K output on tap.
- Aspect ratios
- 1:1 16:9 9:16 4:3 3:4 3:2 2:3
- Reference images
- 1
- Cheapest run
- from 4.5 credits
Clips
Seedance 2.5ByteDanceUp to 30-second take
Thirty reference images, ten reference videos, ten reference audios, and takes up to thirty seconds.
ByteDance's flagship: a full 30 seconds in one pass, with native audio and up to thirty image, ten video, and ten audio references.
- Aspect ratios
- 16:9 4:3 1:1 3:4 9:16 21:9
- Resolutions
- 480p 720p
- Duration
- 4–30 seconds
- Reference images
- 30
- Cheapest run
- from 43 credits
Seedance 2.0ByteDance
The previous Seedance generation, and the only model here that renders 4K.
The previous Seedance: multimodal control from image, video and audio references, and the only one that still reaches 4K.
- Aspect ratios
- 16:9 4:3 1:1 3:4 9:16 21:9
- Resolutions
- 480p 720p 1080p 4K
- Duration
- 4–15 seconds
- Reference images
- 9
- Cheapest run
- from 33.5 credits
Seedance 2.0 MiniByteDanceCheapest
The economical Seedance: same references and ceiling, capped at 720p.
The lowest-cost Seedance: full multimodal reference control and native audio for high-volume drafts and prototypes.
- Aspect ratios
- 16:9 4:3 1:1 3:4 9:16 21:9
- Resolutions
- 480p 720p
- Duration
- 4–15 seconds
- Reference images
- 9
- Cheapest run
- from 17 credits
Seedance 2.0 FastByteDance
Seedance 2.0 tuned for turnaround rather than for price.
The speed-focused Seedance: the same multimodal controls with quicker turnaround for rapid iteration.
- Aspect ratios
- 16:9 4:3 1:1 3:4 9:16 21:9
- Resolutions
- 480p 720p
- Duration
- 4–15 seconds
- Reference images
- 9
- Cheapest run
- from 29.5 credits
Grok Imagine Video 1.5xAI
xAI's clip model. Needs an image to start from, and always generates audio.
xAI's image-to-video preview model: animate a still into a short clip with synchronized music, effects, and ambience.
- Aspect ratios
- 16:9 9:16 1:1 4:3 3:4 3:2 2:3
- Resolutions
- 480p 720p
- Duration
- 1–15 seconds
- Reference images
- Not accepted
- Cheapest run
- from 33.5 credits
Veo 3.1GoogleNative audio
Google's flagship: generates dialogue and sound from the same prompt as the picture.
Google Veo with synchronized, natively generated sound: dialogue and effects baked into the clip.
- Aspect ratios
- 16:9 9:16
- Resolutions
- 720p 1080p
- Duration
- 4, 6, 8 seconds
- Reference images
- 3
- Cheapest run
- from 83.5 credits
Veo 3.1 FastGoogle
Veo without references, at roughly a third of the price. The single-shot tool.
A faster, cheaper Veo 3.1 for rapid drafts before you commit to a full render.
- Aspect ratios
- 16:9 9:16
- Resolutions
- 720p 1080p
- Duration
- 4, 6, 8 seconds
- Reference images
- Not accepted
- Cheapest run
- from 42 credits
Veo 3.1 LiteGoogleCheapest
The volume tier of the Veo family. Audio is compulsory; 1080p only runs at eight seconds.
Google's lowest-cost Veo: native dialogue, effects, and ambience for high-volume drafts and social clips.
- Aspect ratios
- 16:9 9:16
- Resolutions
- 720p 1080p
- Duration
- 4, 6, 8 seconds
- Reference images
- Not accepted
- Cheapest run
- from 21 credits
Kling v3 OmniKuaishou
Kuaishou's flagship: seven references, a reference video, and 4K output.
Kling's omni model: strong motion and lip-sync from images plus a reference video, up to 4K.
- Aspect ratios
- 16:9 9:16 1:1
- Resolutions
- 720p 1080p 4K
- Duration
- 3–15 seconds
- Reference images
- 7
- Cheapest run
- from 70 credits
AHappyHorse 1.0AlibabaBest value
Alibaba's clip model. Prompt-only, five aspect ratios, up to fifteen seconds.
Alibaba's image-to-video model: smooth, affordable animation from a single first frame.
- Aspect ratios
- 16:9 9:16 1:1 4:3 3:4
- Resolutions
- 720p 1080p
- Duration
- 3–15 seconds
- Reference images
- Not accepted
- Cheapest run
- from 58.5 credits
RRunway Gen-4.5Runway
Runway's model. 1080p only, five or ten seconds, six aspect ratios.
Runway Gen-4.5: cinematic camera moves and consistent motion at a flat per-second rate.
- Aspect ratios
- 16:9 4:3 1:1 3:4 9:16 21:9
- Resolutions
- 1080p
- Duration
- 5, 10 seconds
- Reference images
- Not accepted
- Cheapest run
- from 50 credits
P-VideoPruna AICheapest
Pruna's general clip model. Cheap, flexible durations, one reference audio.
Pruna's quick all-rounder for text-to-video, image-to-video, or clips conditioned by your own audio.
- Aspect ratios
- 16:9 9:16 4:3 3:4 3:2 2:3 1:1
- Resolutions
- 720p 1080p
- Duration
- 1–10 seconds
- Reference images
- Not accepted
- Cheapest run
- from 8.5 credits
FLUX 3Black Forest Labs
Black Forest Labs' clip model. Picture and sound from one plain-language prompt, for up to twenty seconds.
Black Forest Labs' multimodal model: one prompt makes the picture and its sound, from nothing at all or from a frame you hand it.
- Aspect ratios
- 16:9 4:3 1:1 3:4 9:16 21:9
- Resolutions
- 720p 1080p
- Duration
- 5–20 seconds
- Reference images
- Not accepted
- Cheapest run
- from 71 credits
MMiniMax H3 MaxMiniMax
fal's own retuning of MiniMax H3. Fifteen seconds, native audio, and almost nothing to configure.
MiniMax's H3, retuned by fal: a prompt, an optional Start and End frame, and up to fifteen seconds that arrive with their own sound.
- Aspect ratios
- 16:9 4:3 1:1 3:4 9:16 21:9
- Resolutions
- 480p 768p
- Duration
- 5–15 seconds
- Reference images
- Not accepted
- Cheapest run
- from 21 credits
P-Video AvatarPruna AILip sync
A photo plus an audio track becomes a talking clip. The prompt is optional.
Turn one portrait into a presenter: attach a voice-over and the face speaks it, with natural mouth and expression.
- Aspect ratios
- 16:9 4:3 1:1 3:4 9:16 21:9
- Resolutions
- 720p 1080p
- Duration
- 1–15 seconds
- Reference images
- 1
- Cheapest run
- from 10.5 credits
P-Video AnimatePruna AI
Takes a still and a driving video, and moves the still the way the video moves.
Recast a shot: your subject performs the motion and audio of a source video. Built for spinning one asset into many UGC variations.
- Aspect ratios
- 16:9 4:3 1:1 3:4 9:16 21:9
- Resolutions
- 720p 1080p
- Duration
- 1–15 seconds
- Reference images
- 1
- Cheapest run
- from 12.5 credits
P-Video ReplacePruna AI
Swaps the subject inside existing footage for one of your own.
Swap the people in an existing shot while keeping its motion, timing, camera, scene, and optional source audio intact.
- Aspect ratios
- 16:9 4:3 1:1 3:4 9:16 21:9
- Resolutions
- 720p 1080p
- Duration
- 1–15 seconds
- Reference images
- 3
- Cheapest run
- from 12.5 credits
OmniHuman 1.5ByteDance
ByteDance's performance model: a portrait and an audio track, at 1080p with real body motion.
ByteDance's film-grade digital human: micro-expressions, gestures, and camera moves that follow the meaning of the audio, not just its syllables.
- Aspect ratios
- 16:9 4:3 1:1 3:4 9:16 21:9
- Resolutions
- 1080p
- Duration
- 1–15 seconds
- Reference images
- 1
- Cheapest run
- from 67 credits
Kling Avatar v2Kuaishou
A portrait and a voice track become a talking clip — people, animals or cartoons alike, at 720p.
Kling's talking avatar: a portrait and a voice track become a performance — humans, animals, cartoons and stylised characters alike.
- Aspect ratios
- 16:9 4:3 1:1 3:4 9:16 21:9
- Resolutions
- 720p
- Duration
- 1–15 seconds
- Reference images
- 1
- Cheapest run
- from 22 credits
Kling Motion ControlKuaishou
Drives a still image with the motion of a video. Prompt optional, 720p or 1080p.
Kling 3.0's motion transfer: your character performs a source video's action, holding identity, scene, and quality across the take.
- Aspect ratios
- 16:9 4:3 1:1 3:4 9:16 21:9
- Resolutions
- 720p 1080p
- Duration
- 3–15 seconds
- Reference images
- 1
- Cheapest run
- from 29.5 credits
Kling Lip SyncKuaishouCheapest
Six credits: re-syncs a person in existing footage to a new audio track.
Re-syncs the mouth of a clip you already have to new speech: attach a voice-over, or let the prompt be the script. Two to ten seconds.
- Aspect ratios
- 16:9 4:3 1:1 3:4 9:16 21:9
- Resolutions
- 1080p
- Duration
- 2–10 seconds
- Reference images
- Not accepted
- Cheapest run
- from 6 credits
Voice
‖Eleven v3ElevenLabs
The most expressive voice in the catalog, and the only one you direct by writing stage directions into the line.
- Voices
- 26
- Languages
- 29
- Text limit
- 3,000 characters
- Cheapest run
- from 8.5 credits
Gemini 3.1 Flash TTSGoogle
Thirty voices across twenty-five languages, from Google.
- Voices
- 30
- Languages
- 25
- Text limit
- 4,000 characters
- Cheapest run
- from 11 credits
KKokoro 82MKokoro
Thirty-three English-accent voices, priced for volume.
- Voices
- 33
- Text limit
- 4,000 characters
- Cheapest run
- from 0.5 credits
MMiniMax Speech 2.8 HDMiniMax
The directable one: emotion selector, speed control, twenty-four languages.
- Voices
- 17
- Languages
- 24
- Text limit
- 4,000 characters
- Cheapest run
- from 4.5 credits
IInworld TTS 2.0inworld
Four voices, sixteen languages, tuned for realtime delivery.
- Voices
- 4
- Languages
- 16
- Text limit
- 2,000 characters
- Cheapest run
- from 2.5 credits