Companion episode — the honest comparison of local AI video tools, with the Wan bench board on screen.

What Wan 2.2 is

Wan 2.2 is a local image-to-video model that ships in two flavors: a light 5-billion-parameter model and a two-expert 14B (A14B) that splits work across a high-noise and a low-noise pass. It’s the “trending” local name right now — so we did the thing nobody in the threads seems to do and benchmarked it on a single 32GB card, honestly.

The real numbers (our bench)

  • 5B — the practical pick. Fast and it fits: ~42–45s per clip at 480p, peaking roughly 20–33GB of VRAM, mostly stable. If you want a Wan clip on consumer hardware, this is the one.
  • 14B (two-expert) — the quality tier, but heavy. It peaks 31–33GB and wedged on about 40% of our runs on 32GB, at ~270–282s per clip. Great when it lands; frustrating when it doesn’t.
  • The VRAM wall is real. 32GB is right at the edge for 14B. A bigger card or renting one changes the story.

Gotchas worth knowing

  • VAE pairing: the 5B uses the Wan 2.2 VAE; the 14B uses the Wan 2.1 VAE. Mismatch it and the decode errors out.
  • The decode wedge: plain high-res-long decodes (720p × 121 frames) hang the GPU to 0% and lock the app. A tiled VAE decode helps, and 720p is stable at shorter lengths (≤ ~73 frames / 3s). Bench at 480p and report the ceiling honestly.

Where it fits

Wan 2.2 is a generation tool, like LTX-2.3 — text/image into motion. It’s a legitimate option, especially the 5B on a mid card. But it’s roughly two years old in model-time, and in real production we’ve quietly swapped Wan shots for LTX. It’s fine. On our machine, it’s not the pick.

Honest verdict

5B: a genuinely usable, light local video model — start here if you’re on 16–24GB. 14B: the quality tier, but it wants a big card and still wedges on 32GB. Worth knowing, worth benching yourself. Just don’t let a hype thread tell you the 14B “runs great locally” — measure it. (Not sure your card can handle it? See which AI rig you need.)