One graphics card with a model too large spilling off it and streaming to disk, beside another card with two workloads crammed inside colliding

Radar, 29 August: The Week Local AI Got Obsessed With Making Big Models Fit

Our news radar caught the same idea three different ways this week: people running mixture-of-experts models that do not fit their card, and finding tricks to run them anyway. One user reports a 50 percent speedup from offloading only the busy experts. Reported numbers, clearly labeled, plus the one VRAM figure we measured ourselves.

August 29, 2026 · 4 min · Aillex / DIY AI
As an Amazon Associate, this site earns from qualifying purchases. Referral links are always disclosed.