One sentence covers the week: the same capability showed up rented and owned, and the free speed had a price tag we only found by looking.
The one that is ours
We swapped the attention backend under our own presenter renders. SageAttention 1.0.6 took a cam pass from 638 seconds to 462, which is 27.7% faster, and it rendered a different woman from an identical prompt and an identical seed. Not a drift, not a wobble. A different person. We only caught it because we compare frames rather than trust wall-clock. The presenter those frames come from is built with our local lip-sync chain.
SageAttention 2.2.0, built from source, is correct and still about 18% faster, and it is what we run now. Our per-pass fallback to sdpa stays in, because a speedup that changes the output is not a speedup.
That number is ours and only ours: one RTX 5090, our stack, our workflow. Your card is not our card. Run the comparison on your own machine before you believe either half of it.
Rented or owned, in three weeks
Grok Bot (xAI, early beta) and Bot Mode for Hermes Agent (Nous Research, MIT) shipped the same idea inside a month: named agent profiles with their own role, model, memory and skills, that can sign in to your tools and come back with finished work. Between them the same architecture is now available at five price points, from a rented seat to a plugin you install and own.
Nothing about the idea is new to anyone who has already run a team of agents locally. What is new is that the consumer packaging arrived in both lanes at once.
The default we almost blamed on the model
Simon Willison measured 22,276 reasoning tokens to produce 3,223 tokens of output from
qwen3.8:27b, because the model defaults reasoning_effort to extra-high. The same prompt with
reasoning off produced 3,715 tokens in 137 seconds instead of 21 minutes. His recommendation, and
ours: run it on low or off first, then turn it up if you actually need it.
We nearly filed this as “the model over-thinks”. It is a setting. That is his measurement, not ours, and it is the second time this year a local model looked broken and turned out to be configured.
Open weights, and a licence to read first
MiniMax Music 3 landed on Hugging Face: full songs up to five minutes, an 8B global LLM plus a 0.6B local LLM, 32 kHz 16-bit stereo out, and a reported path that fits in 8 GB with offload. Check your card before you plan around that number. Open weights, genuinely.
The licence is the part to read before you build a business on it, and it is one click from the model card. We are not lawyers and this is not advice; we are saying the document exists and it has territory and attribution clauses in it, and open weights is not the same sentence as open use.
The ticker
ComfyUI-Krea2-NAG brings negative prompts back at CFG 1.0 · LTX-2 v1.2.0 ships new checkpoints and Gemma 4 text encoders · a ConvRot quant claims Q8 quality at Q6 size (their claim, unverified here) · and a 252-generation test on what actually keeps an AI film visually consistent.
The lab report
The counter says 1,800. The subscriber gate for monetisation is passed. The other gate is watch hours, it needs 4,000 in a rolling year, and we are pacing closer to 1,500. If you want to help and spend nothing: watch the long ones. The benchmarks and the builds, not the sixty-second clips.
Also new: our own agent now reads the firehose. Some of this brief was filed by it rather than found by hand. It is version one and it is days old, so do not picture anything clever.
The kicker
Somebody compressed a 2.9 MB song to 21 KB with Meta’s EnCodec, printed it as eight QR codes and glued them to a paper cassette with an A side and a B side. You can hold the song. You cannot play it, because those 21 KB are not audio, they are numbers that only mean something to EnCodec’s own 24 kHz decoder.
That is the whole local-first argument in one object. The file is not the thing. The file plus the model is the thing. If you have trained a character or cloned a voice, keep a local copy of the decoder as well as the data. A latent without its model is a very small paperweight.
The DIY AI Brief ships every Monday. This one ran late because it needed a QC round, and we don’t ship without QC. Every number we called ours was measured on one graphics card in a home office; everything else is reported by whoever shipped it and labelled that way on screen. All referral links are disclosed. Corrections? r/aillex.
