One sentence covers the week: open models took the majority of real production traffic, and within days four different companies moved to put a price on that road.
The number the whole week reacts to
Open weights hit 62% of production tokens through Vercel’s AI Gateway, up from 28% in June, and four of the top five models on that gateway are Chinese open weights. That is a reported figure from the people who run the gateway, not ours, and it is the context for every other story this week: when the free road starts carrying most of the traffic, people show up to build toll booths.
Local, but rented
Perplexity’s Portable Computer runs an agent entirely on your own machine, and credit where due, it puts a genuinely good consent screen in front of anything that leaves the box. Then it puts the whole thing behind a subscription. Running on your hardware while renting the software is a real category now, and it is worth knowing which half of it you are paying for. The other half is what this channel exists for.
The six billion dollar answer
Nvidia is licensing Poolside’s technology and hiring its engineers into Nemotron, reportedly around six billion dollars of a combined seven billion commitment, to build American open-weight models. Reported numbers, their announcement. The strategic read is simpler than the price tag: the biggest seller of AI hardware just bet that open weights are where the workloads go.
The licence page is the product page
Wan 3.0 took #1 on Artificial Analysis the same week Alibaba’s own page for it reads “Open Source: No.” Meanwhile the same company shipped Wan 2.2-Animate-2 genuinely open under Apache. And we read the Qwen Community Licence 1.0 properly, out loud, because it is not what most coverage says it is. If you build on a model without reading its licence page, the licence page still applies to you. We keep a plain-language walkthrough at read the model licence, and this chapter is the worked example.
The dial, not the model
IBM’s Granite 4.2 ships a switchable thinking mode with an effort budget you set yourself, under Apache 2.0. After the month we spent proving that a locked thinking default can make a good model look broken, a vendor putting that dial in your hands is the right direction, and we said so.
The lab report
Ours, measured here, one RTX 5090 in a home office: capital letters mean nothing to diffusion models. WRITING A PROMPT IN CAPS does not emphasize anything, and one of our all-caps stage directions got typeset into the scene as literal text. Three prompt-craft findings from our own failed takes this week, each one free to test on your own machine, all shown with the takes that failed.
The ticker, and a song on paper
The ticker runs the smaller stories brisk, and the closer is the week’s strangest artifact: a song compressed to 21 kilobytes and printed on paper as eight QR codes. You cannot play it. That is the point.
Every number, labelled
Every number we call ours was measured here and is still only ours: different stack, different card, different result. Everything else is a reported figure from the people who shipped the thing, and it is labelled that way on screen. Go get your own numbers; the tools are free.
Next brief: Ep. 10, every wall became a policy: Torvalds credits an AI in the kernel, Debian votes, a closed format is re-derived overnight, and the safety layer goes on sale.
