Reference · Updated monthly

The image generation model index.

What each model costs, how large it renders, how many reference images it will accept, how often it refuses ordinary work, and which ones we run in production. Ratings come from two independent leaderboards. The capability columns come from calling the endpoints ourselves.

Models from
Google3OpenAI2ByteDance3Microsoft AI3xAI2Reve2+7 more

This page compares the models themselves. If you are choosing between the platforms built on top of them, that is a different comparison: AI fashion model generators, compared.

23
Models tracked
13
Providers
6
Live in On-Model
2026-08
Last refreshed
Safe default
Nano Banana 2

Leave it selected unless the job gives you a reason not to.

Garment text
GPT Image 2

The only one that reads back a crest or woven label correctly.

Widest coverage
Seedream 4.5

Completes swimwear and lingerie where the others refuse.

Next in the queue
Seedream 5.0 Pro

Native 4K and up to 14 reference images, the most here.

The short versionWhat we run today, plus what is being evaluated next. 16 more models are tracked in the full index below.
ModelUse it forEditingPrice / 1kStatusTry it
GPT Image 2OpenAIThe only engine that renders garment text.1463$211AvailableTry it →
Nano Banana ProGoogleHero shots and premium fabric detail.1389$134AvailableTry it →
Nano Banana 2GoogleThe default. Best quality per second of wait.1385$67.0AvailableTry it →
Seedream 4.5ByteDanceWidest subject coverage, loosest adherence.1301$40.0AvailableTry it →
Grok Imagine Image 2.0xAISecond on the editing board. Reachable.1439Candidate
Seedream 5.0 ProByteDanceA large jump over 4.5, at 4K with 14 references.1393$90.0Candidate
Reve 2.1ReveTop of the editing board, 4K, 8 references.1374$200Candidate

Editing score is arena.ai where available, otherwise Artificial Analysis. Jump to the full index of all 23 models ↓

Where each number comes from

Third-party leaderboards

Elo, list price and release date are pulled from the four boards below. We publish both sources side by side rather than picking a favourite, because they disagree often enough to matter.

Our own measurements

Delivered resolution and refusal behaviour for the engines we run are ours, taken by replaying identical production jobs across them. Marked in the table. Ceilings for models we have not run are the provider's own published limits.

What we will not guess

A model we have never called shows not verified rather than a repeated vendor claim. Opinion is confined to the grey line under each model and is labelled as such.

The indexRefreshed 2026-08-13

Every model worth tracking.

Grouped by provider, ranked within each group by editing score. Editing is the column that matters for production work, because every real job starts from an image you already have.

Google

ModelReleasedText‑to‑imagearena / AAEditingarena / AAPriceper 1k imagesMax outputRef. imagesFilteringIn On‑Model
Nano Banana Pro
Gemini 3 Pro Image
Nov 20251246 / 13011389 / 1242$1344K14mediumAvailable

Reach for it on hero shots and premium fabric, where the extra detail is visible and the extra wait is affordable. It takes the same fourteen reference images as Nano Banana 2, though Google guides towards six or fewer when every reference has to be held at full fidelity. It also rendered packshots front-on when a three-quarter angle was requested, so check the framing on angled work.

Content filtering, observed 2026-08: Behaves like Nano Banana 2; same content boundaries.

Try Nano Banana Pro in On-Model →
Nano Banana 2
Gemini 3.1 Flash Image
Feb 20261264 / 13221385 / 1248$67.04K14mediumAvailable

The safe default, and the one to leave selected unless a job gives you a reason not to. Best all-round garment fidelity per second of wait.

Content filtering, observed 2026-08: Blocks a minority of ordinary apparel work; swimwear and lingerie are the usual trigger.

Try Nano Banana 2 in On-Model →
Nano Banana 2 Lite
Gemini 3.1 Flash Lite Image
Jun 20261251 / 12901314 / 1200$33.61Kmultiplenot measuredWatching

Priced for volume and it keeps the reference-image headroom of the full model, but output is 1K only. That rules it out of anything a customer orders at 2K or 4K, and leaves it a possible fast-preview engine rather than a production one.

OpenAI

ModelReleasedText‑to‑imagearena / AAEditingarena / AAPriceper 1k imagesMax outputRef. imagesFilteringIn On‑Model
GPT Image 2
gpt-image-2
Apr 20261381 / 13721463 / 1257$2114KmultiplehighAvailable

The specialist. If the garment carries a crest, a woven label or printed lettering, this is the one engine that reads back correctly. Its output envelope is unusual: the long edge can reach 3840px, but total pixels are capped at 8.3 megapixels, so a wide 16:9 render is a true 4K while a 3:4 portrait tops out near 2450x3260. Everywhere else it is the slowest option and the most likely to refuse the job outright.

Content filtering, observed 2026-08: The strictest of the engines we run. Refuses a large share of unremarkable fashion work, including children's apparel and swimwear; a swimwear recolor returned a hard content block.

Try GPT Image 2 in On-Model →
GPT Image 1.5
gpt-image-1.5
Dec 20251239 / 13131370 / 1251$133not verifiedmultiplenot measuredWatching

Listed for continuity. Its successor is better on both boards and is the one we ship.

ByteDance

ModelReleasedText‑to‑imagearena / AAEditingarena / AAPriceper 1k imagesMax outputRef. imagesFilteringIn On‑Model
Seedream 5.0 Pro
seedream-5-pro
Jul 20261258 / 12821393 / 1247$90.04K14not measuredCandidate

A large jump over Seedream 4.5 on both boards, rendering natively at 4K and accepting up to 14 reference images, the most of anything on this page. If it inherits 4.5's tolerance for the categories other engines refuse, it replaces it outright rather than joining it.

Seedream 4.5
seedream-4-5
Dec 20251147 / 1301 / 1188$40.04KmultiplelowAvailable

We keep it for coverage, not for adherence. It is the weakest of the four at doing exactly what the prompt said, and it ignores the 1K tier entirely. But swimwear and lingerie are real categories with real catalogues, and this is the engine that ships them.

Content filtering, observed 2026-08: The most permissive engine we run, and the only one that completes swimwear work at 4K where the others refuse.

Try Seedream 4.5 in On-Model →
Seedream 4.0
seedream-4
Sep 20251140 / 12211271 / 1185$30.0not verifiedmultiplenot measuredWatching

Listed for continuity. Cheaper than 4.5 and behind it.

Microsoft AI

ModelReleasedText‑to‑imagearena / AAEditingarena / AAPriceper 1k imagesMax outputRef. imagesFilteringIn On‑Model
MAI Image 2.5
mai-image-2.5
Jun 20261256 / 13081402 / 1255$48.11K1not measuredRuled out

Punches far above its price on both boards, and it is out on capability rather than quality. Editing takes one input image, so it cannot be handed a garment and an identity together, and native 1K output cannot serve a 4K tier.

MAI Image 2.5 Flash
mai-image-2.5-flash
Jun 2026 / 1230 / 1233$20.0not verifiednot verifiednot measuredWatching

The lowest price on this page by a wide margin. It shares the MAI line's 1K ceiling, so the same limitation applies.

MAI Image 2.6
mai-image-2.6
1336 / n/anot verifiednot verifiednot measuredWatching

The one to watch on this page. If it ships with multi-reference support it becomes a serious candidate immediately, because 2.5 already scores where it does with one hand tied.

xAI

ModelReleasedText‑to‑imagearena / AAEditingarena / AAPriceper 1k imagesMax outputRef. imagesFilteringIn On‑Model
Grok Imagine Image 2.0
grok-imagine-image-2
1316 / 1439 / 2K3not measuredCandidate

One of the strongest editing scores anywhere on this page, and it accepts up to three reference images at up to 2K. The three-image limit is the thing to watch: it is enough for a garment plus an identity, and tight for anything more.

Grok Imagine (Image Quality)
grok-imagine-image
Apr 20261228 / 12351362 / 1229$50.02K3not measuredWatching

The previous Grok image tier. Same 2K ceiling and same three-reference limit as 2.0, at a similar price but a lower score on both boards, so 2.0 is the one worth evaluating.

Reve

ModelReleasedText‑to‑imagearena / AAEditingarena / AAPriceper 1k imagesMax outputRef. imagesFilteringIn On‑Model
Reve 2.1
reve-2.1
Jul 20261302 / 13251374 / 1259$2004K8not measuredCandidate

On paper the strongest candidate here. It leads the editing board, renders natively at 4K, and takes up to eight reference images, which is comfortably more than our tools send in one call. Nothing structural stands in the way; it needs a bench run.

Reve 2.0
reve-2.0
Jun 20261270 / 12571358 / $24.0not verifiedmultiplenot measuredWatching

Most of 2.1's quality for a fraction of the price. If Reve earns a place, this is the value option to compare it against.

Meta

ModelReleasedText‑to‑imagearena / AAEditingarena / AAPriceper 1k imagesMax outputRef. imagesFilteringIn On‑Model
Muse Image
muse-image
1282 / 1405 / not verifiednot verifiednot measuredWatching

One of the best editing scores on the arena board, with no generally available API to call. Nothing to evaluate until that changes.

Ideogram

ModelReleasedText‑to‑imagearena / AAEditingarena / AAPriceper 1k imagesMax outputRef. imagesFilteringIn On‑Model
Ideogram 4.0
ideogram-4
Jun 20261204 / 1217n/a$60.02K1not measuredRuled out

Historically the strongest text renderer, which is exactly the weakness GPT Image 2 covers for us. Worth re-testing the day a multi-image edit endpoint exists.

Alibaba

ModelReleasedText‑to‑imagearena / AAEditingarena / AAPriceper 1k imagesMax outputRef. imagesFilteringIn On‑Model
Qwen Image 2.0 Pro
qwen-image-2
Apr 20261191 / 12321303 / $75.0not verifiednot verifiednot measuredWatching

Mid-table on both boards at a mid-table price. Nothing yet that would displace what we run.

Black Forest Labs

ModelReleasedText‑to‑imagearena / AAEditingarena / AAPriceper 1k imagesMax outputRef. imagesFilteringIn On‑Model
FLUX.2 [max]
flux-2
Dec 20251162 / 12291262 / 1201$70.0not verifiedmultiplenot measuredWatching

The most widely self-hosted family on this page, which matters if you run your own infrastructure. On hosted quality it sits below what we ship.

Sourceful

ModelReleasedText‑to‑imagearena / AAEditingarena / AAPriceper 1k imagesMax outputRef. imagesFilteringIn On‑Model
Riverflow 2.0
riverflow-2
Feb 2026 / 1275 / 1286$150not verifiednot verifiednot measuredWatching

A useful reminder to read more than one board: it leads the Artificial Analysis editing ranking and does not appear on arena at all, so there is no second opinion to check it against.

Luma AI

ModelReleasedText‑to‑imagearena / AAEditingarena / AAPriceper 1k imagesMax outputRef. imagesFilteringIn On‑Model
Luma UNI 1 Max
uni-1
May 20261188 / 12221334 / 1219$100not verifiednot verifiednot measuredWatching

Consistent mid-table on both boards. No fashion-specific reason to move it up the queue yet.

PiktID

ModelReleasedText‑to‑imagearena / AAEditingarena / AAPriceper 1k imagesMax outputRef. imagesFilteringIn On‑Model
Onda
onda
n/an/a4Kmultiplenot measuredAvailable

Not comparable to anything else on this page, and deliberately so. It does not take a prompt at all: it drives a mask-based workflow that swaps the person while leaving the garment untouched. That is the opposite trade from a general engine, which regenerates everything and hopes the garment survives. Slower, and it can leave artifacts.

Try Onda in On-Model →
Orbita
orbita
n/an/a1Knonenot measuredAvailable

Built for one job: inventing a consistent model identity from a written description, with no source photograph involved. 1K only, which is all an identity reference needs.

Try Orbita in On-Model →

Two ratings per column: arena.ai first, then Artificial Analysis. They are separate populations of voters and separate scales, so compare a model against others in the same column, never one board against the other. A dash means the model is not listed on that board. marks a figure we measured ourselves.

A model is not a productWhere On-Model sits

The engine is raw material. Not the factory.

Every model on this page is a generic API call. It knows nothing about garments, colourways, SKUs, ghost mannequins, or a catalogue of four thousand images that all have to look like they were shot on the same afternoon.

What a model gives you

One image, from one prompt, with no memory of the last one. Ask it twice and you get two different lighting setups, two different models, two different interpretations of the same jacket.

What production actually needs

The same identity across 300 SKUs. A garment that survives untouched. A colourway that matches the swatch. Batch throughput, approval states, retries, and a bill that makes sense at the end of the month.

What sits in between

The prompts, the reference plumbing, the masks, the segmentation, the post-process equalisation, the QC pass. That layer is the product. The engine underneath it is swappable, which is exactly why we keep this page.

If you are comparing the platforms built on top of these engines rather than the engines themselves, that is a different question and we wrote it up separately: AI fashion model generators, compared.

In production6 engines live

What we run, and where.

Every tool exposes an engine picker. Auto is the default on all of them and resolves to the fastest capable engine, with a fallback if the chosen one refuses the job.

ToolSelectable engines
Flat-to-Model
AutoNano Banana 2Nano Banana ProSeedream 4.5GPT Image 2
Create Packshot
AutoNano Banana 2Nano Banana ProSeedream 4.5GPT Image 2
Garment Recolor
AutoNano Banana 2Nano Banana ProSeedream 4.5GPT Image 2
Detail Repair
AutoNano Banana 2Nano Banana ProSeedream 4.5GPT Image 2
Image Playground
AutoNano Banana 2Nano Banana ProGPT Image 2Seedream 4.5
Model Swap
AutoOndaNano Banana 2
Create Identity
AutoNano Banana 2Nano Banana ProSeedream 4.5Orbita

Switching engines is free

We bill by output size, not by engine. A render costs the same whichever model produces it, so you can pick the one that suits the job rather than the one that suits the budget.

Two engines are ours

Onda and Orbita have no leaderboard entry because they are not general-purpose models. Onda swaps the person in a photograph while leaving the garment physically untouched, which is the opposite trade from a general engine that regenerates the whole frame. Orbita invents a model identity from a written description alone.

Model by jobFor fashion and product imagery

Which model for which job.

A leaderboard tells you which image a stranger preferred. This is what each engine is actually good at once a garment is involved, and every one of them is a dropdown away inside On-Model.

Nano Banana 2 for fashion imagery

The safe default, and the one to leave selected unless a job gives you a reason not to. Best all-round garment fidelity per second of wait.

Try Nano Banana 2 in On-Model →7 tools · same price as every other engine

Nano Banana Pro for fashion imagery

Reach for it on hero shots and premium fabric, where the extra detail is visible and the extra wait is affordable. It takes the same fourteen reference images as Nano Banana 2, though Google guides towards six or fewer when every reference has to be held at full fidelity. It also rendered packshots front-on when a three-quarter angle was requested, so check the framing on angled work.

Try Nano Banana Pro in On-Model →6 tools · same price as every other engine

GPT Image 2 for fashion imagery

The specialist. If the garment carries a crest, a woven label or printed lettering, this is the one engine that reads back correctly. Its output envelope is unusual: the long edge can reach 3840px, but total pixels are capped at 8.3 megapixels, so a wide 16:9 render is a true 4K while a 3:4 portrait tops out near 2450x3260. Everywhere else it is the slowest option and the most likely to refuse the job outright.

Try GPT Image 2 in On-Model →5 tools · same price as every other engine

Seedream 4.5 for fashion imagery

We keep it for coverage, not for adherence. It is the weakest of the four at doing exactly what the prompt said, and it ignores the 1K tier entirely. But swimwear and lingerie are real categories with real catalogues, and this is the engine that ships them.

Try Seedream 4.5 in On-Model →6 tools · same price as every other engine

Onda and Orbita are ours and work differently: Onda swaps the person in a photograph while leaving the garment physically untouched, and Orbita builds a model identity from a written description. Both are in the full index above.

What the leaderboards missMeasured August 2026

Three findings you will not find on a board.

Leaderboards measure which image a stranger prefers at a glance. They do not measure whether the logo on a jersey reads correctly, or whether the resolution you paid for is the resolution you got. We replay identical production jobs across engines to answer that.

01

Only one engine renders legible garment text

On a football kit whose crest reads FEDERAZIONE ITALIANA GIUOCO CALCIO, GPT Image 2 reproduced the lettering correctly in every sample. Nano Banana 2, Nano Banana Pro and Seedream 4.5 produced letter-shaped noise in every sample. At full-body framing all four failed, so legibility tracks the pixels the text occupies, not the tier you ordered.

02

The engine with fewer pixels won that test

GPT Image 2 got the crest right at roughly 8 megapixels while Nano Banana delivered 17 and got it wrong. More resolution is not more legibility. It is also worth checking what an engine actually returns rather than what it accepts, which is why the maximum-output column above is measured rather than quoted.

03

Content filtering is a capability, not a footnote

The strictest engine we run declines a meaningful share of unremarkable catalogue work, including children's apparel. The most permissive one is the weakest at following instructions. Swimwear and lingerie are real categories with real catalogues, so coverage and obedience trade against each other and you need both engines available.

The barWhy strong models get ruled out

A high score is not enough.

Several models near the top of both boards are marked ruled out here. That is not a quality judgement. It is a structural one, and these are the four things that decide it.

More than one reference image

Our tools send a garment, an identity and a pose reference in the same call, so a model that accepts one image cannot be asked to hold a jacket and a face at the same time, however well it scores. The ceilings vary more than you would expect: one image at the low end, fourteen at the high end.

A resolution ceiling that reaches our tiers

Customers buy a 4K output. A model that tops out at 1K or 2K cannot serve that, however well it scores, and offering it anyway would mean charging for pixels that never arrive.

A generally available API

Preview access and a leaderboard entry are not an integration. Several of the highest-scoring models on this page have no endpoint anyone outside the vendor can call.

Filtering that tolerates ordinary work

An engine that refuses swimwear, children's apparel or lingerie is not usable as a sole engine for a fashion catalogue. It can still earn a place as a specialist alongside a more permissive one.

QuestionsUpdated monthly

Asked, answered.

Which image generation model is the best?

There is no single answer, which is the reason this index exists. On the leaderboards GPT Image 2 leads text-to-image and the editing boards disagree with each other. On real production work the ranking changes again: the engine that renders legible lettering on a garment is not the engine that handles the widest range of subject matter, and neither is the fastest. Pick per job, not once.

How is Elo calculated, and how much should I trust it?

Both arena.ai and Artificial Analysis run blind pairwise comparisons and convert the votes into an Elo rating, the same system used for chess. It is a good measure of which output a person prefers at a glance, and a poor measure of whether a garment survived unchanged. We publish both boards side by side because they routinely disagree, and a model that only appears on one has no second opinion behind it.

Why do some models show no maximum resolution?

Because we have not called the endpoint ourselves. Capability columns are marked by how we know them: measured on our own bench, read off the provider's live input schema, or unverified. Rather than repeat a vendor claim we have not tested, we leave it blank.

Why are strong models missing from On-Model?

Almost always because they accept a single reference image. Our tools send a garment, an identity and a pose reference in one call, so an endpoint that takes one image cannot do the job at all, however well it scores. The other common reasons are a resolution ceiling below the tier a customer paid for, and no generally available API.

Does choosing a different engine cost more credits?

No. On-Model bills by output size, not by engine, so switching between them is free. That is deliberate: it means you can pick the engine that suits the job rather than the one that suits your budget.

How often is this page updated?

Monthly. Elo, price and release dates are pulled directly from the two leaderboards; capabilities come from our own testing and are dated where they matter. The refresh date at the top of the page is the real one, not a build timestamp.

Use all of them. Pick per job.

Every engine marked available on this page is one dropdown away inside On-Model, at the same price per image. Start with five free generations.