This week open-source AI got enormous — and it barely mattered. The whole story of the week is one contrast: giant open weights nobody can run, versus a tiny one that fits in your pocket.
The 2.4-trillion flex (that you can’t run)
Alibaba previewed Qwen 3.8 Max on July 19 — 2.4 trillion parameters, multimodal, framed as “second only to Fable 5.” No published benchmarks, no model card, and open weights coming “soon” with no date. Days earlier, Moonshot dropped Kimi K3 — 2.8 trillion, weights promised July 27. Two “open” giants in a week. The catch is obvious: 2.4T parameters don’t fit on your rig, your block, or anything short of a server rack. In an open-weight arms race, the announcement itself becomes the weapon — treat every trillion-parameter headline as a claim, not a fact, until the weights and benchmarks actually land.
The 657-megabyte counterpoint (that you can)
Meanwhile a developer fine-tuned OpenBMB’s MiniCPM5-1B on Claude Fable 5’s reasoning traces: a thinking model with tool calls and 128K context that runs in Ollama or LM Studio — smallest quant 657 MB, small enough for a phone. Not a weight distillation, and a 1B won’t rival the giants — but that was never the point. A piece of frontier-grade reasoning, free and offline, in your pocket. That’s what open source was supposed to mean.
Fable 5 won’t die
Anthropic pushed Fable 5’s included access July 7 → 12 → 19, then published a post titled “Redeploying Claude Fable 5.” The model that built this channel keeps refusing to sunset — and now a fragment of it lives inside that free 657 MB model. The letters survive.
The rest of the board
- GLM-5.2 (Z.ai) — the open champ that fits: 744B MoE, tops the real usage leaderboards, runs on serious real hardware.
- LTX-2 landed native in ComfyUI — open 4K, video + synced audio in one pass, up to 3× faster / 60% less VRAM with NVIDIA NVFP4. If you animate images locally, this is your upgrade.
- Ollama v0.32 turned the local runner into an agent (chat, code, web search, delegate) — plus OpenCode for full local coding agents.
- New coding models sized to actually run: Laguna XS 2.1 (Poolside) and Kimi K2.7 Code (Moonshot).
- The practical board: DeepSeek V4 (price), Qwen 3.6 (license), and Gemma 4 12B now beats last year’s 27B flagship at under half the memory.
- At the frontier, the closed labs kept pace (Sonnet 5, GPT-5.6, Grok 4.5). Closed still leads on raw capability; open trails by weeks, not years — and open is the one that’s free, private, and yours.
The DIY AI Brief ships every Monday. Start building at askaillex.com, and bring the week’s questions to r/aillex.
