AN INDEPENDENT PLATFORM FOR CONTEMPORARY ART
ART VALUE LAB
ArtX Digital VaultART VALUE LAB · ART × AI × KNOWLEDGE

Best Performing Models

7 models
PHASE
Leaderboard · All phasesPairwise comparison · Bradley–Terry Elo
170014501200
Muse Image
ChatGPT Images 2.5 Flare
ChatGPT Images 2.0
Nano Banana Pro
FLUX.2 [max]
Grok Imagine 1.0
Krea 2 Medium
Hover, focus, or tap a bar for its score and 95% confidence interval.

General preference · Prompt adherence · Usability · Visual aesthetics

GPT-6 AstraMODEL NOTE

Strong mockups still need a delivery check.

Contra Labs reports stronger results for GPT-6 Astra at the mockup stage, after a direction is established. Ideation, hero design, and text readability still need human judgment. Model preference and complete, working user flows deserve separate checks.

Read the source analysis
MOCKUP · ELO
GPT-6 Astra
1662
GPT-5.6 Sol
1552
Claude Fable 5.1
1489
Claude Opus 5
1489
Muse Spark 1.1
1308
208 retained comparisons; Bradley–Terry Elo. Bars start at 1,300; baseline 1,500.
RankModelElo rating95% CI
1Muse ImageMeta16261605–1647
2ChatGPT Images 2.5 FlareOpenAI16231597–1648
3ChatGPT Images 2.0OpenAI15511531–1571
4Nano Banana ProGoogle15461524–1568
5FLUX.2 [max]BlackForestLabs14571437–1477
6Grok Imagine 1.0xAI13911368–1414
7Krea 2 MediumKrea13061283–1329
METHODS & STANDARDSHow are these models scored? Methods & limitations

Contra Labs Research · Redrawn from source data, not a live ranking or measure of artistic value.

Latest research

Read complete research, from methods and figures to findings and limitations.

39 studies

39 results

Web DesignField Note

How model performance shifted across four website design evals

GPT 5.6 Sol led the first three evals and every loose brief. Once the work was staged across ideation, mockup, and refinement, the Claude models moved ahead.

Read article
GPT win rate vs Fable
Eval 1
59%
Eval 2
72%
Eval 3
70%
HCB
41%
Head-to-head general preference · GPT led three evals, then fell to 41% in HCB
Web DesignBattle

Claude models lead landing page design, but no model wins every design stage

Six models built landing pages for three products across ideation, mockup, and refinement. Opus and Fable finished first and second overall, but a different model led each stage of the work.

Read article
Win rate
Claude Opus 5
59.4%
Claude Fable 5
57.1%
GPT-5.6 Sol
52.1%
Kimi K3
50.7%
Gemini 3.6 Flash
45.5%
Muse Spark 1.1
35.2%
Share of head-to-head comparisons won across 3,240 pairwise decisions · August 2026
VideoBattle

Ad videos: six models across ideation, mockup, and refinement

The Human Creativity Benchmark moves to video: six models, three client campaigns, three phases each, judged head to head by working creatives. Seedance 2.0 won the set, and no model held the product together once it started moving.

Read article
RESEARCH IN NUMBERS
68.3%Seedance 2.0 overall win rate
3,240Pairwise judgments
6 × 3 × 3Models × campaigns × phases
Head-to-head judgments by professional creatives · August 2026
Web DesignField Note

What stands between Opus 5 and client-ready pages: layout and readability

Nine designers annotated 20 landing pages one at a time, 738 notes in all. On Opus 5's pages the notes clustered in layout and readability, while brand fit and originality drew the fewest flags in the set.

Read article
RESEARCH IN NUMBERS
738Annotations across 20 pages
2 in 3Rated Opus 5 issues that needed rework
20 × 9Pages × designers
Solo page annotations, no matchups · August 2026
ImageBattle

Six image models made real ads. Each broke differently.

We scaled the Human Creativity Benchmark into a standing benchmark, starting with ad images: six models, three client campaigns, three phases each, judged head to head by professional creatives. Meta Muse Image won the set, and every model showed a signature failure.

Read article
RESEARCH IN NUMBERS
68.9%Muse Image overall win rate
70.1%ChatGPT Images 2.0 win rate at ideation
6 × 3 × 3Models × campaigns × phases
Head-to-head judgments by professional creatives · August 2026
Web DesignBattle

Claude Opus 5 nails the words and the mood. The finish is what holds it back.

Across 600 blind comparisons and 400 write ups, designers liked the copy more than the rest of the page. Contrast, length and a few broken sections were what held it back.

Read article
Win rate
GPT 5.6 Sol
65.3%
Kimi K3
61.7%
Claude Fable 5
41.7%
Claude Opus 5
31.3%
Share of head-to-head matchups won across 600 blind comparisons · July 2026
Web DesignProfile

Five designers, five rounds: Google Stitch proves its strength in ideation.

Five expert product designers ran five Stitch iterations each on the same dashboard brief. The first prompts delivered wireframes and mockups in minutes; five rounds of refinement never raised the fidelity.

Read article
RESEARCH IN NUMBERS
2 + 3Wireframes and mockup stages after five iterations
42 → 44Issues tagged, V1 → V5
5 × 5Designers × iterations
Think-aloud Stitch sessions · one shared brief · July 2026
Web DesignBattle

Kimi K3 is a real rival to GPT 5.6 Sol on landing pages

Kimi K3 arrived ranked first on Arena's Frontend Code Arena. Across 480 blind comparisons by 8 designers, it finished level with GPT 5.6 Sol on landing pages, and edged ahead on the detailed briefs.

Read article
Win rate
GPT 5.6 Sol
65%
Kimi K3
63.3%
Gemini 3.5 Flash
38.8%
Claude Fable 5
32.9%
Share of head-to-head matchups won across 480 blind comparisons · July 2026
Web DesignField Note

Where four AI models break when they build a landing page

8 designers annotated 40 pages from GPT 5.6 Sol, Claude Fable 5, Grok 4.5, and Muse Spark 1.1, marking 754 failure points across 8 tags. Every model broke differently.

Read article
RESEARCH IN NUMBERS
754Failure points marked
40Pages · 4 frontier models
8Designers annotating
Per-page failure annotation · 5 briefs · July 2026
ImageBattle

GPT Image 2 won 41.9% of logo tournaments, and still only half its logos were client-ready

10 brand designers judged GPT Image 2, Nano Banana Pro, MAI Image 2.5, and Meta Muse across 12 logo briefs. The winner is clear, and still only half its logos cleared the client bar.

Read article
Tournament win rate
ChatGPT Images 2.0
41.9%
Nano Banana Pro
22.6%
MAI Image 2.5
20.6%
Meta Muse
17.4%
Share of logo tournaments won · 4 image models · 12 briefs · July 2026
ImageField Note

Pretty isn’t the same as right: One image model runs away with brief fidelity

Five designers, four criteria, 6,400 blind pairwise ratings. Nano Banana 2 takes first on every brief-fidelity axis while the aesthetics standings invert behind it.

Read article
RESEARCH IN NUMBERS
4/4Fidelity criteria swept
67%Top typography win rate
52%Best rival ceiling
Nano Banana 2 vs the field · 6,400 blind ratings · July 2026
ImageField Note

No model owns “aesthetics”: What 8,000 designer ratings tell us about taste in image models

Five designers, five criteria, 8,000 blind pairwise ratings. Three of the four frontier models land within a few points of each other, and the best model changes depending on which dimension of visual quality you care about.

Read article
Pooled win rate
FLUX.2 [max]
54%
Nano Banana 2
52%
GPT Image 1.5
50%
Seedream 5.0 Lite
44%
Pooled win rate across 5 aesthetic criteria · 1,600 ratings each · July 2026
ImageBattle

Reve 2.1 trailed Seedream 5.0 Pro by 2 points on wins, then finished last on Elo.

Four-model image battle across 10 briefs spanning portrait, environment, product, and lifestyle work, judged blind by working creative professionals.

Read article
Tournament wins
Seedream 5.0 Pro
32.6%
Reve 2.1
30.2%
MAI Image 2.5
20.9%
ChatGPT Images 2.0
20%
Share of tournaments won · Seedream 5.0 Pro leads Reve 2.1 by 2.4 pp
Web DesignBattle

Sol has taste. Fable takes direction.

GPT 5.6 Sol, Claude Fable 5, Grok 4.5, and Muse Spark 1.1 on the same ten landing page briefs, judged blind by 9 working designers as live interactive pages.

Read article
Win rate
GPT 5.6 Sol
63.3%
Claude Fable 5
49.3%
Grok 4.5
48.1%
Muse Spark 1.1
39.3%
Share of 540 pairwise matchups won · 10 landing page briefs · July 2026. Split by brief structure the ranking inverts: Fable leads structured briefs at 1569 Elo to Sol's 1455.
ImageBattle

Seedream 5.0 Pro trails ChatGPT Images 2.0 by 8 points on wins, takes photorealism.

Four-model image battle across 10 briefs spanning the capabilities ByteDance advertises. Seedream won photorealistic generation at 35.7% and finished first or second in 57% of its tournaments.

Read article
Tournament wins
ChatGPT Images 2.0
35.9%
Nano Banana Pro
28.6%
Seedream 5.0 Pro
28%
Flux 2
10.7%
ChatGPT Images 2.0 leads Seedream 5.0 Pro by 7.9 pp
ImageBattle

Meta Muse trailed ChatGPT Images 2.0 by 5 points on wins, then beat it on Elo.

Four-model style-transfer battle across 10 briefs. Muse took 29.5% of tournament wins and was the only model top-two on both wins and Elo.

Read article
Tournament wins
GPT Images 2.0
34.1%
Meta Muse
29.5%
Nano Banana Pro
23.9%
Flux 2
19.3%
Share of 88 tournaments won · ChatGPT Images 2.0 leads Meta Muse by 4.6 pp
Web DesignBattle

Write Fable a design spec and it wins 9 times out of 10

Claude Fable 5 vs Claude Opus 4.8 on 5 real landing page and portfolio briefs, built in Claude Code, judged blind by 9 working designers.

Read article
RESEARCH IN NUMBERS
88.9%Best brief win rate
11.1%Worst brief win rate
51.1%Fable overall vs Opus
Fable 5 vs Opus 4.8 · 90 blind matchups · July 2026
VideoBattle

Frontier AI video models are nearly tied. None of them nail physics yet.

A blind head-to-head of Seedance 2.0, Grok Imagine, Veo 3.1, and Adobe Firefly Video, judged by 12 professional video editors across 10 prompts.

Read article
Prompts won
Seedance 2.0
4
Grok Imagine
3
Veo 3.1
3
Adobe Firefly Video
0
Prompts won out of 10 · 4 frontier video models · June 2026. Seedance 2.0 also led on average score, by +0.07 over Grok Imagine on a 5-point scale.
ImageField Note

Introducing Design Crit: we taught AI to judge design like a designer.

Ten professional designers ranked four frontier image models across nine dimensions of real design work. The models can make the work. Nothing on the market could reliably judge it, until we trained on the right data.

Read article
Agreement with panel
Human designer
74.1%
Trained on Design Crit
61.1%
Best off-the-shelf
54.3%
Chance
50%
Agreement with the five-designer majority · best off-the-shelf judge = HPSv2.1 · June 2026
ImageBattle

Ideogram v4 won 47.9% of typography matchups.

10 designers, 4 models, 240 images. Spelling is solved. Typographic craft and client-readiness are where Ideogram v4 pulls away.

Read article
Typography 1st-place rate
Ideogram v4
47.9%
Gemini 3.1 Flash
30%
FLUX.2 [max]
15.5%
Grok Imagine 1.0
15%
Share of 1st-place finishes on typography prompts · 4 image models · 20 prompts × 10 reviewers · rounds 1 + 2
ImageProfile

Gemini reliably edits, but can it keep the rest of the image still?

11 production-style sessions. Gemini made the edit 73% of the time, kept the rest of the image still 64%, held both in 55%.

Read article
RESEARCH IN NUMBERS
8 / 11Sessions that passed the edit-isolation test
7 / 11Sessions that passed the pose-lock test
6 / 11Sessions that passed both controllability tests
Gemini controllability checks · 11 sessions · two tests per session (edit isolation + pose lock)
Web DesignBattle

Cursor took 60% of head-to-heads. Claude Code took 63% of client meetings.

Four coding tools, 24 outputs, five working designers. The tool designers preferred to look at and the tool they'd put their name on turned out to be different.

Read article
Client-ready
Claude Code
63%
Antigravity
53%
Cursor
47%
Codex
30%
“Would you present this to a client?” · 4 coding tools · 24 outputs · May 2026
ImageProfile

Gemini hit production-ready 24% of the time. One prompt pattern explains why.

10 participants, 29 scored deliverables. The prompts that landed treated Gemini like a creative brief for a specific asset.

Read article
RESEARCH IN NUMBERS
7 / 29Deliverables Production Ready
16 / 29Deliverables at Client V1+
2 / 10Participants Production Ready on all 3 deliverables
Gemini (Nano Banana Pro) production readiness · 10 participants, 29 scored deliverables
ImageProfile

Can Adobe Firefly edit like Photoshop?

8 targeted edits across 4 designer sessions. Firefly cleanly resolved 1, drifted on 5, and missed 2.

Read article
RESEARCH IN NUMBERS
1 / 8Edits cleanly resolved
5 / 8Edits landed partial
2 / 8Edits unresolved
Adobe Firefly Edit · 8 attempts across 4 sessions (Hero + Social per session)
Cross-cuttingProfile

The only prompt that got videos to production-ready in Adobe Firefly.

4 designers, 3 deliverables each. The prompts that landed described physical direction, not aesthetic mood.

Read article
RESEARCH IN NUMBERS
1 / 4Videos production-ready, first pass
2 / 4Social stills production-ready, first pass
20Photoshop mentions across 4 sessions
Adobe Firefly designer evaluation · 4 designers, 3 deliverables each
Web DesignField Note

In Claude Design, your opening prompt decides the ceiling.

5 designers, 5 openings, 1 luxury brief. The first prompt set what each session could reach.

Read article
Specificity score (0–1)
00.250.50.751P1P2P3P4P5
First-prompt specificity by participant, Claude-scored · 5 designers, 1 brief
Web DesignProfile

Claude Design gets you 40%, Figma gets the rest.

5 sessions, 5 designers, 1 real-world client brief. Strong as a starting structure, breaks under precision edits.

Read article
RESEARCH IN NUMBERS
60 → 100%Designers flagging layout & spacing, Edit 1 → Edits 4–5
5 / 5Sessions where layout was the recurring failure mode
≤40%Designer verdict: use it to here, then hand off
Claude Design designer evaluation · 5 designers, 1 real-world client brief
ImageField Note

With ChatGPT Images 2.0, "Text is solved." Typography isn't.

42 sessions, 7 designers. ChatGPT Images 2.0 nails the typographic system, then breaks on the individual characters.

Read article
RESEARCH IN NUMBERS
+3 / +3 / +1Macro themes (hierarchy, brand fit, fonts) net positive
−1 / −3Micro themes (legibility, size & weight) net negative
0Designer mentions of size & weight as a strength
Typography sentiment · 42 sessions, 7 designers, 6 briefs
ImageProfile

ChatGPT Images 2.0 won every head-to-head. Here's where it still breaks.

41 sessions, 7 designers, 6 briefs. GPT Image 2 nails the concept, then breaks at production.

Read article
RESEARCH IN NUMBERS
5 / 41Sessions shipped from GPT alone
33 → 59%Typography "no issues" climb
60-65%Realism plateau, every iteration
ChatGPT Images 2.0 production readiness
ImageBattle

The image-model leaderboard flips by brief.

Four frontier image models, six brand campaigns, ranked blind by working creatives. GPT Image 2 wins the aggregate. Every other model owns a category.

Read article
1st-place rate
ChatGPT Images 2.0
40.7%
Seedream 5.0 Lite
22.3%
FLUX.2 [pro]
22.2%
Gemini 3.1 Flash
14.8%
Share of 1st-place rankings across 6 brand campaigns · 4 frontier image models
ImageBattle

Krea 2 Large is the #2 style-transfer model, closing on GPT Image 2.

Four-model style-transfer evaluation. Krea took #2 on style fidelity, 0.14 points behind GPT Image 2.

Read article
Style Fidelity (avg)
ChatGPT Images 2.0
3.53
Krea 2 Large
3.39
Gemini 3 Pro
2.74
Seedream 5.0 Lite
2.42
Style Fidelity average rating · Krea 2 Large takes #2, 0.14 points behind ChatGPT Images 2.0
ImageBattle

Seedream 5.0 Lite swept the field on product detail shots.

A blind head-to-head against the leading image models from Google, OpenAI, and Black Forest Labs, evaluated by professional creatives.

Read article
Win rate
Seedream 5.0 Lite
63.9%
Gemini 3 Pro
52.8%
GPT Image 1.5
44.4%
FLUX.2 [max]
38.9%
Pairwise win rate · 4 leading image models · March 2026
Cross-cuttingField Note

Creatives keep telling us the same thing about AI: every output looks the same.

12 models, 5 creative domains. One repeated complaint from working evaluators: the work all looks the same.

Read article
CONCEPT MAP
High steerability
Creative partnerFull-spectrum toolUnreliableOpinionated engine
Low best-practiceHigh best-practice
Low steerability
Convergence (best-practice) and divergence (steerability) as orthogonal signals.
VideoProfile

Grok Imagine is the "Polisher" model. Hand off the early rounds, bring it in for refinement.

The biggest phase-over-phase climb of any video model in the study. 3rd at ideation, 1st at refinement.

Read article
Win rate
Ideation
46%
Mockup
44%
Refinement
56%
Grok Imagine win rate · +10pp climb from ideation to refinement
Cross-cuttingField Note

The creative process has 3 phases. AI performs very differently in each.

Ideation, mockup, refinement. AI fits differently at each phase, and the best creatives know where to hand off.

Read article
RESEARCH SNAPSHOT
Ideation
Loose grip
Mockup
Narrowed
Refinement
Firm grip
How tightly creatives hold control across phases · Qualitative
VideoProfile

Veo 3.1 is the "Creative Director" model. Use it early, but hand off before refinement.

61% win rate at ideation. 39% at refinement. The clearest model profile in our video evaluation.

Read article
Win rate
Ideation
61.1%
Mockup
55.6%
Refinement
38.9%
Veo 3.1 win rate · −22.2pp drop from ideation to refinement
Cross-cuttingField Note

Solo creatives are earning more with AI and staying independent.

Higher earning potential, more projects, no new hires. The survey from working independents.

Read article
RESEARCH SNAPSHOT
66%

of independent creatives report higher earning potential since adopting AI26% no · 8% other

Survey · Independent creatives on Contra
Web DesignBattle

We tested 4 AI models with professional web designers. Claude won, but not the way you'd expect.

Claude Opus 4.6, Gemini 3.1 Pro, ChatGPT 5.3 Codex, Qwen 3.5. The winner shifted at every phase.

Read article
Leading win rate
Ideation
79.8%
Mockup
68.9%
Refinement
60%
Per-phase leader shifts · Preview of the Human Creativity Benchmark
Cross-cuttingField Note

AI isn't replacing creative professionals. It's making the best ones better.

Survey of high-earning independent creatives. What they actually do with AI on real client work.

Read article
RESEARCH SNAPSHOT
<25%

of AI output makes it to final deliverablesThe rest is stripped, reworked, or scrapped

Dominant survey response · Independent creatives