Companion episode — the honest comparison of local AI video tools, with the Wan bench board on screen.
What Wan 2.2 is
Wan 2.2 is a local image-to-video model that ships in two flavors: a light 5-billion-parameter model and a two-expert 14B (A14B) that splits work across a high-noise and a low-noise pass. It’s the “trending” local name right now — so we did the thing nobody in the threads seems to do and benchmarked it on a single 32GB card, honestly.
The real numbers (our bench)
- 5B — the practical pick. Fast and it fits: ~42–45s per clip at 480p, peaking roughly 20–33GB of VRAM, mostly stable. If you want a Wan clip on consumer hardware, this is the one.
- 14B (two-expert) — the quality tier, but heavy. It peaks 31–33GB and wedged on about 40% of our runs on 32GB, at ~270–282s per clip. Great when it lands; frustrating when it doesn’t.
- The VRAM wall is real. 32GB is right at the edge for 14B. A bigger card or renting one changes the story.
Gotchas worth knowing
- VAE pairing: the 5B uses the Wan 2.2 VAE; the 14B uses the Wan 2.1 VAE. Mismatch it and the decode errors out.
- The decode wedge: plain high-res-long decodes (720p × 121 frames) hang the GPU to 0% and lock the app. A tiled VAE decode helps, and 720p is stable at shorter lengths (≤ ~73 frames / 3s). Bench at 480p and report the ceiling honestly.
Where it fits
Wan 2.2 is a generation tool, like LTX-2.3 — text/image into motion. It’s a legitimate option, especially the 5B on a mid card. But it’s roughly two years old in model-time, and in real production we’ve quietly swapped Wan shots for LTX. It’s fine. On our machine, it’s not the pick.
Honest verdict
5B: a genuinely usable, light local video model — start here if you’re on 16–24GB. 14B: the quality tier, but it wants a big card and still wedges on 32GB. Worth knowing, worth benching yourself. Just don’t let a hype thread tell you the 14B “runs great locally” — measure it. (Not sure your card can handle it? See which AI rig you need.)
