CONTRA LABS / Battle
GPT Image 2 won 41.9% of logo tournaments, and still only half its logos were client-ready
Loading the complete research…
Related research
Six image models made real ads. Each broke differently.
We scaled the Human Creativity Benchmark into a standing benchmark, starting with ad images: six models, three client campaigns, three phases each, judged head to head by professional creatives. Meta Muse Image won the set, and every model showed a signature failure.
Contra LabsRead article
Pretty isn’t the same as right: One image model runs away with brief fidelity
Five designers, four criteria, 6,400 blind pairwise ratings. Nano Banana 2 takes first on every brief-fidelity axis while the aesthetics standings invert behind it.
Contra LabsRead article
No model owns “aesthetics”: What 8,000 designer ratings tell us about taste in image models
Five designers, five criteria, 8,000 blind pairwise ratings. Three of the four frontier models land within a few points of each other, and the best model changes depending on which dimension of visual quality you care about.
Contra LabsRead article