A Mozilla analysis landed on the radar from two directions yesterday, and the headline number is a good one: open-weight models are now roughly 4.4 months behind the closed frontier.

The method is more interesting than a leaderboard. Rather than scoring answers, it fits against task horizon data: how long a job, measured in how long a human expert would take, a model can carry without falling over. The best closed model is handling work that takes an expert around twelve hours. The best open-weight models reach that same bar about four and a half months later. Tom’s Hardware has the write up.

Other numbers in the same story are just as striking. On one popular gateway, eight of the top ten models by token volume are open weight. The leading open model trails the leading closed one by three points on a composite intelligence index while costing a fraction as much.

If you run models at home, you probably read that and felt good. We did too, for about a minute.

The model doing the catching up

The open-weight side of that comparison is carried by Kimi K3. We looked its weights up on Hugging Face rather than trusting a summary, and the parameter count is 2.78 trillion.

That is not a model you run. It is not a model you run on a very good gaming PC, or two of them, or the machine you were thinking of building next year. It lives on a rack.

So the 4.4 month gap is real, and it is a gap between two kinds of datacenter. One of them publishes its weights, which matters enormously for researchers, for companies who want to self host, and for the general health of the field. It does not mean the thing on your desk is four months behind Anthropic.

We are not being sour about this. Open weights at that scale are how the small stuff eventually exists, because distillations, quantisations and smaller siblings all come downstream of it. But “open” and “yours” are two different words and the coverage keeps merging them.

The gap that actually applies to you

There is a second gap, and nobody publishes a number for it: the distance between the frontier and what fits on hardware you own.

That one is wider. It is also closing, and by completely different means. Not by anyone training a bigger model, but by the unglamorous work of making existing ones fit.

In the last two weeks alone, on this radar: a video model with generated sound running on integrated graphics, the same class of model doing lip sync inside 8 GB of VRAM, a brand new music model quantised from a serious card to 8 or 9 GB within a day of release, and somebody training a small image model from scratch on one GPU in three and a half days. This morning there are people moving key value cache into system RAM to buy context length back on cards that should not have it.

None of that will make a headline, because the number does not move. The model is the same model. It just arrives on hardware that could not hold it last month.

What we would actually tell you

The honest version has three parts and only one of them is about benchmarks.

The frontier is further ahead of your desktop than four months, and anyone quoting that figure at you about a local model is comparing the wrong things. The report’s own authors are more careful than the coverage: the analysis notes that open models may be fitted more aggressively to public benchmarks, and that closed labs do not always release their strongest systems, so four months is a public estimate rather than a settled fact.

For the work most people actually do, the gap is not the binding constraint anyway. We run a whole video channel on one consumer card, and nothing we have been blocked on this year was blocked by model intelligence. It was VRAM, or a licence, or a dependency, or a full disk.

And the direction of travel is good in the way that matters to you specifically. The frontier moving is somebody else’s news. Something you could not run in August running in September is your news, and that has happened four or five times since we started writing these notes.

A daily note from our news radar. The week gets the full treatment in the DIY AI Brief every Monday.