The written companion to EP40. In July we benched twelve generators on one RTX 5090 (that guide is here). This is the rematch: twenty on the boards, the same rules, and a harder test. Every number below is ours, measured on one card, one generation per cell. Treat it as a starting point, not a spec sheet.

The rules, and the two new ones

Same intent to every model. One generation each, same seed, no re-rolls, no picking the pretty one. The prompt is printed on every board so you score adherence yourself. Local models ran free on our RTX 5090 in ComfyUI 0.34.3; cloud models ran through Civitai’s nodes and we logged the exact Buzz each image cost.

New rule one: the specifics are invented. July’s logo said “Ember and Oak”, and another channel ran the exact same logo, which tells you the name lives in the training data. A name a model has seen a thousand times tests recall, not rendering. So every prompt this round names things that do not exist: a shop called Quill and Tallow, a cafe called The Tallow Room, a flooded limestone quarry with one red canoe.

New rule two: longer text is the real test. Text is a ladder now. A logo with three words. An infographic with three labels and small captions. A chalkboard with a heading, five menu lines with prices and a closing sentence, about thirty words. And because one roll proves nothing about spelling, every local model ran four extra seeds on each rung, Google’s Gemma read every cell back, and we scored what it read against the words we asked for.

Callouts on the boards are for value (what it costs you) and adherence (did it draw what the prompt asked for). We do not pick a best image. That is taste, and it is yours.

Who ran

Local, free on the cardSteps / guidanceWeights on diskWarm seconds per image
Z-Image Turbo8 / 1.020.4 GB4.0
Z-Image Base50 / 4.020.4 GB30.1
Krea 2 Turbo8 / 1.018.4 GB6.0
Krea 2 Raw52 / 3.531.5 GB58.2
CyberRealistic Krea 28 / 1.031.5 GB10.0
Krea 2 Dark Realism8 / 1.018.4 GB4.1
LiquidMix Krea 2 Turbo8 / 1.018.4 GB4.0
Qwen-Image-2512 (fp8)50 / 2.529.8 GB88.3
Qwen-Image-2512 Lightning 4-step4 / 1.031.5 GB10.0
Flux.1-dev (fp8)20 / guidance 3.517.2 GB8.0
Anima 2B24 / 4.05.4 GB6.0
Illustrious28 / 6.07.4 GB6.0
SDXL DreamShaper Lightning8 / 2.06.9 GB4.0
Cloud, through CivitaiSeconds (incl. queue)Buzz per image
GPT Image 2146194
Nano Banana 214104
Nano Banana Pro14160
Google Imagen 4did not rundid not run
Seedream1660
Flux 21480
Grok Image1826
Qwen-Image-2.0 (via fal)1646

Warm means the model was already loaded; the first image of each run carries the load time and is excluded. The “VRAM” your monitor shows during these runs is the allocator’s cache, 28 to 30 GB even for a 6-billion-parameter model, so the number that matters is weights on disk. Imagen 4 returned “workflow failed, no failure details” on every attempt, including prompt-only probes: a hosted model you cannot debug is a hosted model you cannot debug.

Qwen Image 3.0 is closed. It went API-only in August 2026 with no weights, no licence and no report, so it is a reported model, not a tested one. What we could run is Qwen-Image-2512, the newest open release (20B, Apache 2.0), plus its Lightning LoRA, and Qwen-Image-2.0 through Civitai’s fal node.

The eight boards

Click any board for full size. Each cell carries the arm, seconds, weights or Buzz, and the prompt style it was given.

Landscape

A flooded limestone quarry at dawn under low fog, a rusted crane on the far ledge, one red canoe on the still water.

Landscape board

Scenes no longer separate the field. Eighteen or nineteen of twenty land the quarry, the fog and the crane. The adherence check is the small stuff: SDXL gave two canoes, Anima a forest of cranes standing in the water, Illustrious a beautiful canyon with no crane. Value: Z-Image Turbo and LiquidMix Krea 2 Turbo at four seconds, free.

Anime

An old lighthouse keeper in a yellow raincoat feeds seagulls from a lantern gallery in a storm, cel shaded, crisp lineart.

Anime board

Anima 2B again: 5.4 GB, six seconds, free, and this is what it was built for. Krea 2 Raw and Flux shrank the keeper to a speck on the tower; SDXL put him on the rocks instead of the gallery; Nano Banana 2 turned the frame into a movie poster with a title in Japanese nobody asked for.

Digital art, the July callback

A night market built inside the ribcage of a beached whale skeleton, paper lanterns between the ribs, concept art.

Digital art board

In July the big models built a library where the small ones gave us bookshelves. Same split: the Krea 2 family, Qwen, GPT Image 2, Nano Banana, Seedream and Flux 2 built the ribcage as architecture with the market inside it; Z-Image Turbo, Anima, Illustrious and SDXL drew a live whale over a market. Flux built a skull and put the market in its jaws. The “grandeur gap” is still there, but it is no longer local versus cloud: Krea 2 runs on this card and built the cathedral.

Portrait, the one bee

Studio portrait of a middle-aged woman beekeeper in a canvas apron holding a honeycomb frame, one bee on her cheek, window light, 85mm.

Portrait board

Eighteen of twenty read as that woman. The tell is the bee: Anima turned the frame into a hoop crusted with bees; Illustrious drew a cartoon bee and a younger woman, July’s finding exactly; SDXL lost the bee. When every model can paint the scene, the test is one small specific the prompt insisted on.

Action

A skateboarder mid-air over a flooded underpass, spray from the front wheels, a green traffic light reflected in the water, low angle.

Action board

Nineteen of twenty land the jump, the spray and the green light. Anima doubled the skater, Illustrious went abstract, and GPT Image 2, at 194 Buzz and about two and a half minutes for this cell, cropped the rider’s head off. Value: SDXL DreamShaper Lightning, four seconds on 6.9 GB.

Logo, rung one of the ladder

A clean modern logo for a candle and stationery shop called QUILL AND TALLOW, bold legible lettering, flat vector, muted palette.

Logo board

On this seed Z-Image Turbo wrote “Tallow” with one L, Anima wrote “QUL”, Illustrious wrote something that is not English on a book, SDXL managed the word twice, differently. The rate table below is the real result.

Infographic, rung two

A poster titled LOCAL AI IN 2026 with three labelled bars: VRAM 8 GB, 16 GB, 32 GB, and a small readable caption under each.

Infographic board

Most local arms clean. Anima grew a fifth bar. Flux blurred the whole flat graphic, which it did not do on its logo re-run, so “Flux fp8 and flat art” is an open question at two rolls, not a verdict. Grok added a word; Flux 2 drew the 16 and 32 bars the same height; Nano Banana 2 decorated the poster until the bars were the smallest thing on it.

The chalkboard, the top rung

A cafe chalkboard reading THE TALLOW ROOM, then five menu lines with prices, then the sentence “Ask about the Thursday candle class.”

Longtext board

Krea 2 Turbo, Raw, Dark Realism and LiquidMix rendered all seven lines. CyberRealistic wrote OYE for RYE. Z-Image Turbo misspelled QUINCE and CARDAMOM; Z-Image Base dropped “class”. Qwen-Image-2512, base and Lightning, both wrote QUINQE. Flux held three lines and collapsed. Anima, Illustrious and SDXL produced writing-shaped noise. Every cloud model held the whole board.

The text ladder as a rate

Four extra seeds per local arm per rung, every cell transcribed by Gemma 4 26B and scored as the share of expected words recovered. Gemma is itself a model, so a wrong reading is possible in both directions.

ArmLogoInfographicLongtext
Krea 2 Turbo, Dark Realism, LiquidMix1.001.001.00
CyberRealistic Krea 21.000.991.00
Krea 2 Raw1.001.00 (0.89 exact lines)0.98
Qwen-Image-25121.001.000.98 (0.86 exact)
Qwen-2512 Lightning 4-step0.921.000.98
Z-Image Base1.001.000.99
Z-Image Turbo0.83 (0.50 exact)1.000.94 (0.79 exact)
Flux.1-dev0.670.610.93 (0.71 exact)
Anima 2B0.920.170.71
Illustrious, SDXLat or under 0.25at or under 0.09unreadable
Every cloud model (one seed)1.001.001.00 (Qwen-Image-2.0: 0.97)

Three readings. The Krea 2 family and Qwen-Image-2512 sit at the top of the ladder on this card. Z-Image Base beats Z-Image Turbo on text at every rung, at 30 seconds against 4, which gives July’s “Base is not automatically better” a text exception. And the two families on the older architecture, Illustrious and SDXL, fail every rung: old versus new, not local versus cloud, still.

Did rewriting prompts per model change anything?

Every prompt was rewritten per family: dense hundred-word prose for Z-Image, short natural sentences for Krea 2, Flux and Qwen, tags for the booru families, because copying one prompt between different text encoders is supposed to be unfair. So we tested it. Every model got the wrong style on purpose on the portrait, the logo and the chalkboard.

Longtext ablation

On the modern models it changed almost nothing. Krea 2 spelled the chalkboard either way; Z-Image spelled the logo either way; the beekeeper kept her bee and her apron on every big model regardless of prose style. The only families that moved were Anima, Illustrious and SDXL, and they fail both ways. We rewrote the prompts anyway; it did not change the ranking. (Portrait, logo.)

Native 2K

Krea 2 is trained to 2048 pixels square, so the chalkboard and the quarry ran at native 2K on the four fast models, one seed.

Longtext at 2K

Krea 2 Turbo held every word and every price at 26 seconds. Z-Image Turbo lost prices and whole lines. Qwen Lightning wrote QUINQE and CARDUMOM. Flux collapsed. 2K costs roughly three to five times the 1K seconds on this card. (The quarry at 2K.)

What it costs

The cost board

Cloud, per image: Grok 26 Buzz, Qwen-Image-2.0 46, Seedream 60, Flux 2 80, Nano Banana 2 104, Nano Banana Pro 160, GPT Image 2 194 (and the slowest). Local: nothing, every seed, on a card you already own. The argument from July is stronger now, because the local models at the top of the text ladder are the free ones. The cloud gives you one good image for a price; the card gives you a hundred tries for free, and on the board that mattered most, the free ones held.

Traps and caveats

  • The VRAM column is the allocator. nvidia-smi peaks at 28 to 30 GB on nearly every arm because the allocator caches; weights on disk is the actionable number.
  • Gemma read the text cells. The rate table is a model reading a model. Spot-check against the boards.
  • n=1 per bench cell, n=4 on the text rungs. A starting point, not a spec sheet. Any cell can be re-rolled and re-prompted, and locally that is free.
  • Flux.1-dev fp8 and flat art is unresolved at two rolls.
  • Imagen 4 would not run through Civitai on 2026-09-06, this round’s Ernie.

The prompts

Each genre’s prompt in all three renderings (dense, natural, tags), exactly as sent: landscape, anime, digital art, portrait, action, logo, infographic, longtext. Ask in the comments if you want a specific cell at full resolution.

The lab report on this bench, and the plate lesson it cost, is in the DIY AI Brief, Ep. 10. July’s guide, the original bench: 12 AI image generators, honestly compared. Prompting Krea 2 and its LoRAs: Krea 2 prompting and LoRAs.