Two unrelated posts turned up this morning, from different people using different models, and they are the same discovery.

The first: draw a red circle on your reference image and MiniMax H3 puts the scene there. Their words, and worth quoting because it is the part that surprised them: a red circle drawn in the water makes the scene happen in the water. They also say it is not perfect and some detail goes missing, which we appreciate more than a clean claim.

The second: a LoRA for Flux 2 Klein 9B that controls where eyes look, in any image, in any style, by placing a red dot where the gaze should go. The same author previously did one for sun direction, which is the same idea again: a position is easier to point at than to describe.

Neither of those is ours and we have not reproduced either. Treat them as reports.

Why this rings true to us

Because our own most useful result last week was the same principle in a completely different channel.

We spent a day testing H3 and the finding that mattered was about audio. Handing the model our presenter’s voice as a reference works for five or six seconds, after which it starts re speaking the words in its own performance and drifts. Handing the same audio to a node that pins it from frame zero held it for a full fifteen seconds, with correlation at 0.95 or better in every two second window we measured.

Same audio. Same model. The difference was which input we used.

There was a second one in the same day that we did not connect at the time. On a still portrait, we asked one checkpoint for no cuts and a static camera, in the prompt, in plain English. It cut to a close up at about three seconds anyway, and changing the seed only changed when. Switching to the checkpoint whose job is holding a reference held the shot for fifteen seconds, every time, with no instruction about cuts at all.

We told it what we wanted and it ignored us. We picked the right input and it complied without being asked.

The useful version of this

There is a habit in this hobby of treating the text box as the whole interface, and of responding to a model that does not obey by writing more words at it. Longer prompt. More adjectives. Negative prompt full of things you do not want. We do it too.

But most of these systems have several inputs, and they are not equal. A reference image, a mask, a guide track, a first and last frame, a control image, the choice of checkpoint itself. Words are the loosest of them, and they get interpreted. A circle drawn on a picture is not interpreted, it is a coordinate.

So when a model is ignoring an instruction, the question worth asking before you rewrite the sentence is whether the thing you are asking for has a channel of its own. Position usually does. Timing usually does. Identity usually does. Mood mostly does not, which is why prompt wording still matters for tone and atmosphere and increasingly little for placement.

What we would test next

Whether the circle survives, and how far it goes. Two subjects with two circles, for a start. Whether colour matters or any high contrast mark will do. Whether it survives on a last frame as well as a first, and whether it degrades at the resolutions people actually render at.

We have not run any of that, so we are not going to pretend to know. It is on our list, and if it holds up it changes how we build reference plates, because we currently spend real effort describing a position in words that a shape could specify exactly.

If you want the measured version of the audio finding rather than the summary, it is in our H3 write up, along with the rest of a day’s testing on one card.

A daily note from our news radar. The week gets the full treatment in the DIY AI Brief every Monday.