
Radar, 29 August: The Week Local AI Got Obsessed With Making Big Models Fit
Our news radar caught the same idea three different ways this week: people running mixture-of-experts models that do not fit their card, and finding tricks to run them anyway. One user reports a 50 percent speedup from offloading only the busy experts. Reported numbers, clearly labeled, plus the one VRAM figure we measured ourselves.


