The arena
Every week we run the same prompts through every model we can dispatch, publish the outputs side by side, and rank what came back.
We do not make any of these models. justpictur.ing pays every one of them per generation, so when a model wins here the only thing that changes for us is where the work gets routed. That is the whole reason we can run this.
Image models
| # | Model | Rating | Move | Won | Median cost |
|---|---|---|---|---|---|
| 1 | GPT Image 2OpenAI | 1591 | new | 2 of 4 | 11 credits |
| 2 | Nano Banana 2Google | 1579 | new | 2 of 4 | 6 credits |
| 3 | FLUX.2 ProBlack Forest Labs | 1443 | new | 0 of 4 | 2.5 credits |
| 4 | Grok Imagine ImagexAI | 1443 | new | 0 of 4 | 2 credits |
| 5 | Seedream 5 LiteByteDance | 1443 | new | 0 of 4 | 3 credits |
Cost is what the wallet actually charged that generation, not an estimate. There is deliberately no speed column: our generation records carry a start time and no finish time, so any duration we printed would be measuring our own polling loop rather than the model.
Video models
The video board opens with the first week that runs the video suite. Clips cost real money per second, so that suite runs against a contender pool rather than the whole catalog — and it has not run yet.
This week
Week 36, 2026 · August 31, 2026
Arena week 36: every model can spell now
All five models rendered FERMÉ LE LUNDI correctly, accent and all. Two years ago that would have been the whole story; this week it was the least interesting result on the board.
Read the full week
How this is scored
Every model runs the same prompt, at the same aspect ratio, on its own default settings. Defaults are what a new account gets, and tuning each prompt per model would measure our patience rather than the model.
Ratings are Elo, starting at 1500 with K=32. A matchup is one event in which the winner beat every other entrant at once, so all of its rating changes are computed against the standings as they stood when it began and applied together — the order the entrants happen to be listed in cannot change the result. The whole history is replayed from scratch on every build, so a rating is never stored and can never drift from the results it summarises.
Judging is editorial: we pick the winner, and we say why on each matchup. It is not a community vote and this page will not pretend otherwise. Every prompt is published in full and every output is here at full size, so the work is checkable — which is the only kind of ranking worth linking to.
Run the winner yourself
Every model on this board is one dropdown in the same studio. Copy any prompt on this page, pick a different model, and see what you get.
Start picturing