Yesterday morning the author of a MiniMax H3 prompt builder posted an update saying they had cut identity bleed on a feature called RefMods. This morning somebody else posted that they were blown away by them, having previously tried character sheets, image to video, and training actual LoRAs to solve the same problem.

Their description is the interesting part. They called it, in scare quotes, a way to make a “LoRa” of a character from a handful of untagged images, producing a safetensors file you keep forever.

The scare quotes are doing real work, because no training happens at any point.

What the file actually contains

We read the repository rather than the enthusiasm. RefMods take pictures, video clips or audio, run them through H3’s existing VAE, and save the resulting latents to a small .safetensors in models/refmods. That is the whole creation step. An encode and a write.

At generation time a RefMod Text Encode node presents those stored latents during tokenization, numbers them so you can cite them in a prompt as <Video 1> or <Audio 1>, and gives each one a strength slider.

So it ends up looking exactly like a LoRA. Small reusable file, drops into a models folder, loaded at generation, has a strength. It is also the opposite of a LoRA in the one way that decides what it can do. A LoRA changes the model’s weights. A RefMod changes what you hand the unchanged model at inference. It is a cached encoding, not learning.

Why that distinction decides your expectations

If nothing is trained, several things follow without anyone having to test them.

The speed and the lack of captions are free. There is no training run to sit through, no caption pass and no epochs to pick between. No overfitting either, because there was never a fit. That matches both reports exactly.

It also cannot teach the model anything it could not already represent. Training moves the weights toward something the base model was not doing. Encoding does not. Everything a RefMod can give you is something H3 could already produce if you had handed it the right reference at the right moment, which is a large set, and is not an unlimited one.

And the limits are budget limits rather than quality limits. The repo documents a 12 slot stack, a token budget that blocks creation when exceeded, a cap of 8 pictures the encoder sees by default, and a VAE that stores two frames for a clip of 17 frames or fewer and five more for each additional 17. Audio keeps only the first few seconds you nominate. Those are the numbers that will decide whether it works for your shot, and none of them is about how good the model is.

The audio part is the one we find most interesting. One file can carry a face and a voice, which is a different unit of reuse from anything we currently keep.

Where we stand

We have not run it. Everything above is the repository’s own documentation plus two users’ reports, and we are saying that once rather than hedging every sentence.

It matters to us because we solved this problem the expensive way. Our presenter has a trained H3 identity LoRA behind her, built from a photo set and voice cuts, and the training run and the selection around it were a day of work. If a handful of untagged images gets within sight of that for a single episode, the pipeline changes.

What we will measure when it goes on the bench is the thing neither post covers. Not whether the face looks right in one clip, which both authors have shown convincingly, but whether it holds across many separate renders at different framings over an episode, which is the case where our trained version earns its cost. Those are different questions and the second one is the one that costs money to get wrong.

The repository is MIT licensed with about 250 stars, written by one person, and had a commit today.

A daily note from our news radar. The week gets the full treatment in the DIY AI Brief every Monday.