One graphics card with a model too large spilling off it and streaming to disk, beside another card with two workloads crammed inside colliding

Radar, 29 August: The Week Local AI Got Obsessed With Making Big Models Fit

Our news radar caught the same idea three different ways this week: people running mixture-of-experts models that do not fit their card, and finding tricks to run them anyway. One user reports a 50 percent speedup from offloading only the busy experts. Reported numbers, clearly labeled, plus the one VRAM figure we measured ourselves.

August 29, 2026 · 4 min · Aillex / DIY AI
The Local Model Gauntlet: gemma4 vs qwen3.8 vs muse-glimmer

The Local Model Gauntlet: gemma4 vs qwen3.8 vs muse-glimmer

Three open-weight models, one RTX 5090, and the same jobs: a real bug, playable games, tool calling, long documents. Measured numbers, two defaults that were quietly breaking the results, and the finding we published wrong and had to correct.

August 24, 2026 · 11 min · Aillex / DIY AI
Give Your Local AI Eyes: Screen and Camera Vision with Ollama

Give Your Local AI Eyes: Screen and Camera Vision with Ollama

Your local model can probably already see, how to check in 60 seconds, three working vision patterns (screenshot describe, eyes for a text-only brain, live companion vision), and the three gotchas that cost us real debugging hours.

July 10, 2026 · 6 min · Aillex / DIY AI
The Everyday Local-LLM Toolkit: PDFs, Summaries, Study Guides, and More

The Everyday Local-LLM Toolkit: PDFs, Summaries, Study Guides, and More

One local model, six daily superpowers, chat with your documents, digest long videos, summarize voice memos, build study guides, prep interviews, draft anything, all offline and private.

July 9, 2026 · 3 min · Aillex / DIY AI
As an Amazon Associate, this site earns from qualifying purchases. Referral links are always disclosed.