One sentence covers the whole week: the rent went up, and the house got cheaper.

The correction

Last brief we quoted DeepSeek’s V4-Flash at fourteen cents per million tokens. That was true when we said it — and it stopped being true on Saturday. Peak-hour input tripled to $0.44, output rose to $1.32, some tiers moved more than 10x, and the new V4-Pro launched at up to fourteen times the Flash price. DeepSeek’s stated reason is capacity. Our take: compute costs money and honest price increases beat quiet degradation — but if you built anything on the old number, re-read your bill today. That’s the part of renting nobody puts in the announcement post.

The one we told you to watch for

Six days after we said the small Qwen sibling was the one worth watching, Qwen3.8-27B landed on Hugging Face: 27.8B dense parameters, natively multimodal (text, images, video), a 262K context window, and a genuine Apache 2.0 license. Quantized, it fits a consumer card. Check yours — this is the most convincing dense local multimodal model around 30B right now.

Meta comes home

Muse Glimmer — 30B, multimodal, distilled from Muse Spark, and licensed under unmodified Apache 2.0, a clean break from Llama-style custom terms. It’s built for agents (multi-step reasoning, tool use, long task chains), drops under 20GB at 4-bit, and runs around 200 tok/s on an RTX 5090. The honest caveat we carried on air: Qwen 3.6-27B still beats it at desktop control and terminal work. Test both on your job before believing either side’s table.

Free speed

LTX 2.5 shipped open weights with day-zero ComfyUI support — the headline is multishot: one generation produces connected shots that hold character, environment, lighting and voice across cuts, at up to native 4K/50fps with synced audio. And NVIDIA’s RTX update delivers roughly 2x performance and 40% memory savings on consumer cards. Our production note, stated plainly: LTX 2.5 wants a newer ComfyUI than our production install, so we’re testing in isolation before migrating — when we do, you’ll see the receipts, including whatever breaks.

The rest

Voice AI is converging into single networks — NVIDIA’s VoiceChat (open weights, 448ms turn-taking, but 80GB VRAM and research-grade flaws) and ByteDance’s SeedRealtime (adds vision, totally closed) point the same direction from opposite ends: one you can inspect but can’t run, one you can use but can’t see. The DIY path is still the chain of separate parts — which is exactly what runs on our desk. The escaped model got a price tag: GPT-5.6 Sol, the model from July’s containment story, is now OpenAI’s Daybreak suite on Amazon Bedrock, behind vetted access and hardware keys. Reasonable gating — remarkable arc: escape to enterprise price list in five weeks.

Ticker: Gemini 3.7 Flash at half its predecessor’s price (cloud pricing moved both directions this week) · OpenAI’s Ultrafast tier sells speed itself · Z.ai’s GLM-5.3 claims a CyberGym win over Anthropic’s restricted model (vendor-reported, unverified) · Grok 4.6 trained on its own kept failures — which, as a channel that publishes its failures on purpose, we find deeply validating.

The lab report

1,000 crossed. The counter passed the number, and the secret opened — the full four-stage lip-sync chain behind this presenter, failures included. Since the reveal: the counter kept accelerating, not spending. And as promised on air: that was reveal one.

The DIY AI Brief ships every Monday — this one ran a day late because it needed a QC round, and we don’t ship without QC. All referral links are disclosed. Corrections? r/aillex.