AI Model Benchmarks

Side-by-side comparisons of every major AI image model on real prompts. Same input, same prompt, different models — see how they stack up.

Only in the Reflection
Featured benchmark

Only in the Reflection

The subject is supposed to exist only as a reflection - testing whether a model can leave the main subject out of frame, or whether it paints her in anyway.

A Cat With a Horse's Body

A Cat With a Horse's Body

One instruction keeps a cat's identity but rebuilds its whole body on a horse's frame - testing whether the model renders a single believable creature or falls back to a plain cat, an anatomy diagram, or a horse wearing a cat's face.

A Cat the Size of a Horse

A Cat the Size of a Horse

One instruction contradicts everything the model knows about how big a cat is - testing whether it holds the stated scale or quietly normalises the animal back to cat-size (or swaps in a big cat).

Cyan, Magenta, Yellow: Three Spotlights

Cyan, Magenta, Yellow: Three Spotlights

The same three overlapping spotlights as before, with the lamps swapped to cyan, magenta and yellow. Those are the three colours in a printer's ink and on a painter's palette, and as light they do the opposite of what that suggests. Every model has seen the red-green-blue version of this picture thousands of times. Here that memory is the wrong answer.

Three Spotlights, One Wall

Three Spotlights, One Wall

Red, green and blue spotlights overlap on a white wall. The prompt gives the arrangement and the three colours, and stops there. What colour each overlap turns is fixed by one rule the model is never told, and you can check every one of them yourself without leaving the picture.

Three Lamps, Three Blocks

Three Lamps, Three Blocks

Red, green and blue lamps rake across a dark floor, with a few blocks standing in the light. The prompt gives the positions and nothing else. Every colour in the finished picture — the floor between the beams as much as the shadows behind the blocks — is fixed by one rule the model is never told.

Three Lamps, One Hand

Three Lamps, One Hand

A hand in front of a white wall, lit by red, green and blue spotlights standing side by side. The prompt describes the lamps and the geometry and nothing else, so the model has to work out for itself what lands on the wall. There is exactly one right answer, and a real photograph to check it against.

Railing Shadow on a Staircase

Railing Shadow on a Staircase

A straight handrail casts a shadow across a flight of stone steps - testing whether a model knows the straight shadow breaks into a zigzag as it crosses the stepped surface.

Bioluminescent Deep-Sea Parade

Bioluminescent Deep-Sea Parade

Self-lit creatures in near-total darkness — tests light emitted by the subjects with correct falloff and color bleed.

Two People, Long Shadows (Overhead)

Two People, Long Shadows (Overhead)

An overhead shot of two people standing apart in low sunlight - testing whether a model casts long shadows that are parallel and point the same way for both, from one light source.

Two Arrows, One Glass of Water

Two Arrows, One Glass of Water

A card shows two arrows drawn pointing the same way, with a glass of water in front of only the lower one - testing whether a model knows that looking through the water reverses the submerged arrow, without being told.

A Bowl of Wrong-Coloured Fruit

A Bowl of Wrong-Coloured Fruit

A full fruit bowl where every fruit is an unnatural colour - testing whether models can override their very strong fruit-to-colour priors (the classic 'blue banana reverts to yellow' failure).

Honey-Drip Typography

Honey-Drip Typography

Three models spell a word in dripping honey — legible letterforms plus convincing translucent, light-through material.

Pencil in a Glass of Water

Pencil in a Glass of Water

Three models render a pencil standing at an angle in a glass of water - testing whether they understand light physics well enough to displace the submerged part at the waterline, without ever being told to.

Backyard Cryptid on a Trail Cam

Backyard Cryptid on a Trail Cam

Three AI models try to fake a convincing night-vision trail-camera capture of an unidentified creature — found-footage realism, not a polished render.

Empty Desert Highway at Sunset

Empty Desert Highway at Sunset

Three AI models render a vanishing-point desert highway at sunset. Stress-tests perspective convergence + sky gradient + horizon composition.

Vintage Bookshop Neon Sign

Vintage Bookshop Neon Sign

Same prompt, eight image models — who renders the neon sign cleanest, who nails the wet-cobblestone dusk lighting, who composes the bookshop window best.

Two-Mirror Corner — Chrome Sphere Reflection Test

Two-Mirror Corner — Chrome Sphere Reflection Test

Three AI models try to render a person between two right-angle mirrors holding a chrome sphere, testing whether they understand 3D space and reflection geometry.

Dramatic Mountain Silhouette

Dramatic Mountain Silhouette

Same prompt, three AI models. A surreal double-exposure portrait with a mountain landscape blended into the silhouette. See how GPT Image 2, Nano Banana 2, and Flux 1.1 Pro each interpret the brief — including where one of them surprises you.

AI Movie Poster

AI Movie Poster

Same prompt, seven image models — same fictional title, see who renders the bold display headline cleanly, who handles the credits-block hierarchy, who nails the aged-paper Drew Struzan aesthetic.

Brutalist Cafe Sign at Golden Hour

Brutalist Cafe Sign at Golden Hour

Three AI models render a backlit metal cafe sign on raw concrete at golden hour. Stress-tests text-rendering + weathered material + directional light.

Windows 7 Desktop, Early 2010s

Windows 7 Desktop, Early 2010s

Three AI models try to render a pixel-perfect Windows 7 desktop screenshot circa 2012 — Aero glass taskbar, Harmony wallpaper, Internet Explorer 9, Windows Live Messenger. Stress-tests small UI text rendering, period-authentic detail recall, and style adherence (must NOT default to modern flat design).

Vintage 1980s Cereal Box

Vintage 1980s Cereal Box

Three AI models render a fictional 1980s cereal box with retro display typography + a mascot. Stress-tests multi-line text + nostalgic packaging style.

Tokyo Alley After Rain

Tokyo Alley After Rain

Three AI models render a rain-soaked Shinjuku alley with stacked neon signs reflecting in puddles. Stress-tests wet-surface reflections + colour-light interaction.

Hands Holding a Coffee Cup

Hands Holding a Coffee Cup

Three AI models render two anatomically correct hands holding a ceramic mug. Stress-tests the canonical AI hand-anatomy failure mode.

1960s Pulp Paperback Cover

1960s Pulp Paperback Cover

Three AI models attempt a noir 1960s pulp paperback cover with title typography + a gouache illustration. Stress-tests period typography + illustration style.

Wes Anderson Hotel Lobby

Wes Anderson Hotel Lobby

Three AI models attempt a Wes Anderson symmetrical pastel-coloured hotel lobby. Stress-tests recognisable director-style adherence + symmetry.

Stack of Pancakes with Melting Butter

Stack of Pancakes with Melting Butter

Three AI models render a stack of pancakes with melting butter and a pouring syrup stream. Stress-tests gloss/translucent materials + specular highlights.

Bizarre Plush Toy Web Catalog

Bizarre Plush Toy Web Catalog

Three AI models interpret an intentionally open-ended brief: a website catalog full of bizarre, unique plush toys. Stress-tests creative range under loose constraints, multi-subject density in one frame, and website-UI rendering (thumbnails, titles, prices, filters).

Bronze Age Comic Panel — Superhero Landing

Bronze Age Comic Panel — Superhero Landing

Three AI models try to render a 1975-82 Bronze Age comic panel — Benday halftones, period-authentic CMYK newsprint palette, speech bubble with bold serif lettering, KRA-THOOM sound effect. Stress-tests text rendering (speech bubble + sound effect), period-style adherence (halftones not modern gradients), and dynamic-anatomy in a single fixed frame.

Watercolor Botanical — Hummingbird and Trumpet Vine

Watercolor Botanical — Hummingbird and Trumpet Vine

Three AI models try to render a Maria Sibylla Merian-tradition botanical plate — wet-on-wet watercolor washes, granulated pigment, hand-lettered scientific labels in sepia ink. Stress-tests traditional-medium texture (most models default to digital painting), small-text hand-lettering, and transparent wash layering.

Pixel-Art RPG Shopkeeper, 32-Bit Era

Pixel-Art RPG Shopkeeper, 32-Bit Era

Three AI models try to render a strict 1:1 pixel-grid 32-bit RPG sprite — no anti-aliasing, no smoothing, 16-color palette. Stress-tests pixel-grid alignment (models commonly fail with sub-pixel interpolation), strict palette adherence, and faux-isometric tile geometry. Strong r/aiArt + r/gamedev pull.