Monday, 17:17 UTC, on r/StableDiffusion: you don’t need the character swap LoRA for MiniMax H3. The reference checkpoint already swaps characters natively, stylised or realistic, without any adapter.
Tuesday, 06:29 UTC, same subreddit: you do need the character swap LoRA for MiniMax H3.
Both posts contain real testing. Both authors are describing what happened on their machine accurately. They are answering different questions, and you cannot tell from either headline.
What each one was actually running
The second post opens, below the title, with an ellipsis: "… if you want to use the hybrid model." The condition is right there, one line under a headline that could not carry it.
MiniMax ships two checkpoints. FL2VA takes a first and last frame and makes the move between them. REF2VA takes reference pictures, video or audio and keeps what you gave it. The first poster is on REF2VA, where character swapping is a native capability, so from where they sit the LoRA is an optional boost. They say as much: it helps hold the reference on fast motion, and they posted a side by side showing it.
The second poster is on neither. They are running a community hybrid, a weight level merge of the two official checkpoints made by one person on Hugging Face, whose own model card explains the trade: you get FL2VA’s picture quality with REF2VA’s reference conditioning, and some abilities take a hit. Video referencing is one of them. The LoRA brings it back.
So one person is saying a native feature does not need an adapter, and the other is saying an adapter restores something a merge broke. Those are compatible statements about three different sets of weights, and “MiniMax H3” is the name all three go by.
Why this is worth your attention beyond one LoRA
Search Hugging Face for MiniMax H3 and you will find the official repository, the ComfyUI repack, turbo LoRAs from at least three authors, INT8 and NVFP4 and GGUF conversions, pruned variants, upscalers, accelerator LoRA sets, and several independent hybrid merges. The character swap LoRA at the centre of this has just under 6,000 downloads.
Every one of those changes the answer to “does H3 do X”. None of them changes the name.
That is not anybody’s fault and it is the cost of a healthy ecosystem, which we have written about approvingly when strangers filled a gap in a music model in five days. But it does mean a confident sentence about a model is worth roughly nothing without the checkpoint attached, and almost nobody attaches it, because the title box is short and the answer feels general when it is not.
What we can add from our own runs
We publish with H3 and the same fork shows up in our work, from a different direction.
We measured both official checkpoints on one card and wrote the numbers up. Given the same still portrait, FL2VA decided around three seconds that it wanted a close up and cut to one, and no amount of telling it not to made a difference. REF2VA, given that identical picture as a reference, held a single shot for the full fifteen seconds every time. Same model family, same machine, same input, opposite behaviour.
We also trained our own identity LoRA for our presenter, which sounds like the thing the two posts are arguing about and is not. Swapping a character inside one clip is a problem you can often solve with a reference frame. Keeping the same person recognisable across dozens of shots and multiple episodes is a different problem, and that is what a trained LoRA is for. Both are called character consistency.
The version of this you can use
Before you take or give advice about a local video model, say which weights you are on. Checkpoint, any merge, any LoRAs and their strengths, and the step count. It is one line and it converts an argument into two useful reports.
And when two posts disagree, read past the headline before picking a side. In this case the disagreement lasted thirteen hours and dissolves entirely in the first sentence of the second post.
A daily note from our news radar. The week gets the full treatment in the DIY AI Brief every Monday.
