The companion episode — every board on screen, with the prompt printed so you can score adherence yourself.

How we tested

Twelve generators. Six prompts — landscape, anime, digital art, a logo, a portrait, an action shot. The same prompt and the same seed to every model, one generation each, no re-rolls and no picking the pretty one. Local models ran on our own RTX 5090 and cost nothing. Cloud models ran on Civitai’s fleet and cost Buzz, and we recorded the exact price each one charged.

One generation is not a verdict. Any of these can be re-rolled and re-prompted — and locally you can do that as many times as you like for free. That freedom to iterate is itself part of the comparison, and it’s the reason the cheap local models punch above their weight in real work.

The numbers nobody publishes

ModelWhereSpeedVRAMCost per image
Anima 2Blocal~7.5s7 GBfree
SDXL DreamShaperlocal~8s8 GBfree
Illustriouslocal~9s9 GBfree
Z-Image Turbolocal~14s21 GBfree
Flux.1-devlocal~16s24 GBfree
Z-Image Baselocal~37s24 GBfree
Grok Imagecloud~23s26 Buzz
Google Imagen 4cloud~13.5s40 Buzz
GPT Image 2cloud~70s53 Buzz
Seedreamcloud~20s60 Buzz
Flux 2cloud~54s80 Buzz
Nano Banana Procloud~17.5s160 Buzz

That’s a six-times spread between frontier cloud models for the same prompt — and every local model did the same job for nothing.

Three findings worth your time

1. Anima runs on a card you probably own

Seven gigabytes, about seven seconds an image. Most “you need a monster GPU for AI art” advice is out of date — a 2-billion-parameter model holds its own against models that cost real money, especially on anime and stylised work.

2. “AI can’t spell” is out of date — and the split isn’t where you think

Ten of twelve rendered a logo reading EMBER AND OAK correctly. The two that failed are both running locally — and both are older architectures. The newest local models spell fine: Flux clean in 9 seconds, Z-Image in 12, Anima in 6 on 7 GB.

It’s not local versus cloud. It’s old versus new. If a model can’t do text, that’s its generation — not the fact that it’s on your machine.

3. Prompt adherence fails quietly

Our portrait prompt asked for a weathered older fisherman in a worn canvas jacket and knit cap. Eleven models delivered exactly that. One returned a young woman — jacket right, cap right, lighting right, subject completely wrong. Some models are trained so hard toward one kind of image that they’ll quietly overrule your subject, and you’ll only catch it if you’re comparing against the prompt.

Related: our first version of that prompt said “deep skin texture” and half the models stripped the poor man to the waist. Your wording does more work than you think.

Where the cloud genuinely wins

We went in expecting local to hold up everywhere. It didn’t. On a floating library in the clouds, glowing books drifting in the air — a whole invented place rather than an object — the cloud models built an actual library: scale, architecture, painterly depth. Most local models gave us bookshelves, and some went cartoony.

If what’s in your head is grand imaginative scale, a bigger model is more likely to hand you something grandiose. That’s worth paying for when it’s what you need.

Two traps we fell into

  1. “Base” is not automatically better than “Turbo.” We ran Z-Image Base on Turbo’s settings — few steps, almost no guidance — and it came out looking worse and slower than Turbo. Turbo is distilled to look good cheaply; Base only earns its edge with real steps and guidance (60 steps, cfg 5 in our case). Match the settings to the model.
  2. One model wouldn’t run at all. Ernie accepted every job we sent and never returned one — we gave it twenty-five minutes on a single image. Not a verdict on the model, just the honest report.

What we actually use

For everything on this channel: local, because we iterate constantly and don’t want to pay per attempt.

  • Modest card (7–8 GB): start with Anima 2B or SDXL DreamShaper.
  • Big card: Flux.1-dev for prompt adherence.
  • Grand imaginative scale: a big cloud model, because that’s the one thing this bench showed they genuinely do better.

Free and local for the ninety percent you iterate on; pay for the shot that needs a bigger imagination.

Not sure what your machine can run? See which AI rig you need. Want to make video instead? See local AI video tools, honestly compared.