This week open-source AI got enormous — and it barely mattered. The whole story of the week is one contrast: giant open weights nobody can run, versus a tiny one that fits in your pocket.

The 2.4-trillion flex (that you can’t run)

Alibaba previewed Qwen 3.8 Max on July 19 — 2.4 trillion parameters, multimodal, framed as “second only to Fable 5.” No published benchmarks, no model card, and open weights coming “soon” with no date. Days earlier, Moonshot dropped Kimi K32.8 trillion, weights promised July 27. Two “open” giants in a week. The catch is obvious: 2.4T parameters don’t fit on your rig, your block, or anything short of a server rack. In an open-weight arms race, the announcement itself becomes the weapon — treat every trillion-parameter headline as a claim, not a fact, until the weights and benchmarks actually land.

The 657-megabyte counterpoint (that you can)

Meanwhile a developer fine-tuned OpenBMB’s MiniCPM5-1B on Claude Fable 5’s reasoning traces: a thinking model with tool calls and 128K context that runs in Ollama or LM Studio — smallest quant 657 MB, small enough for a phone. Not a weight distillation, and a 1B won’t rival the giants — but that was never the point. A piece of frontier-grade reasoning, free and offline, in your pocket. That’s what open source was supposed to mean.

Fable 5 won’t die

Anthropic pushed Fable 5’s included access July 7 → 12 → 19, then published a post titled “Redeploying Claude Fable 5.” The model that built this channel keeps refusing to sunset — and now a fragment of it lives inside that free 657 MB model. The letters survive.

The rest of the board

  • GLM-5.2 (Z.ai) — the open champ that fits: 744B MoE, tops the real usage leaderboards, runs on serious real hardware.
  • LTX-2 landed native in ComfyUI — open 4K, video + synced audio in one pass, up to 3× faster / 60% less VRAM with NVIDIA NVFP4. If you animate images locally, this is your upgrade.
  • Ollama v0.32 turned the local runner into an agent (chat, code, web search, delegate) — plus OpenCode for full local coding agents.
  • New coding models sized to actually run: Laguna XS 2.1 (Poolside) and Kimi K2.7 Code (Moonshot).
  • The practical board: DeepSeek V4 (price), Qwen 3.6 (license), and Gemma 4 12B now beats last year’s 27B flagship at under half the memory.
  • At the frontier, the closed labs kept pace (Sonnet 5, GPT-5.6, Grok 4.5). Closed still leads on raw capability; open trails by weeks, not years — and open is the one that’s free, private, and yours.

The DIY AI Brief ships every Monday. Start building at askaillex.com, and bring the week’s questions to r/aillex.