Last episode we turned a presenter into an anime one and the rule that came out of it was hand the model a picture. This episode is the photoreal version of that method, and the rule sharpened into something less comfortable:

Whatever is in the picture is coming with it. A reference is not a face. It is everything in the frame, and the model reads all of it as an instruction, including the parts you did not notice were there.

Everything below ran on one RTX 5090 at home, in ComfyUI, with models anybody can download.

The presenter-chain graphs from the episode, as ComfyUI API json (what the /prompt endpoint takes; read it for the node chain and the numbers): ep46_seedhunter_presenter_api.json. Presenter shots, the films and every H3 clip in the episode were generated with MiniMax H3, by MiniMax, used under the MiniMax H3 licence. Nothing generated by H3 was used to train anything.

The bar, and whose work this is built on

On 18 September u/Relevant_Eggplant180 posted Everything Is Melting on r/comfyui: a four-minute photoreal music video, one woman singing a whole song to camera through about forty environments, and when we ran our cut detector over it there is not one hard cut in 4:02. It is a single continuation chain. They built it on two other people’s work, the SeedHunter preview workflow by Fox Fur Essence and Seitanism’s ComfyUI-H3-Motion-Context-MultiRef, with the LBH-123-AI H3 latent upscaler, and released the whole thing as a package with a guide. Go and look at theirs first. Everything here is us taking that mechanism apart and reaching for it with our own presenter in the picture.

The mechanism, in four jobs

The graph looks like a lot. It does four things.

Job one: identity by reference. Instead of starting the video from a still frame (how we made the anime), the model gets reference pictures, and every picture is told in the prompt what it is for. The prompt opens with MiniMax’s own reference grammar:

<Subject 1> is the main performer. <Picture 1> defines facial identity and
<Picture 2> defines full-body appearance, wardrobe, proportions, and
distinguishing details.

followed by a retention_analysis block that says what is fully preserved (face, hair, proportions, wardrobe, colours) and what is free to move (pose, expression, framing, lighting). A third picture, given a job the same way, stages anything else that must appear: a car, a companion, a prop.

Job two: the chain. Each new clip starts from the last 39 frames of the previous clip, held in the model’s latent space rather than re-encoded from a video file. We measured the seam: the mean pixel difference across the overlap is 2 on a 0-255 scale and there is no spike at frame 39. The join is invisible at the pixel level. Where it is not invisible is at the story level, and that is below.

Job three: locked audio. The whole soundtrack (a song, or in our case a voice recording) is loaded once as the master. The workflow cuts the exact section each clip needs from the clip number and the frame count and protects it through both passes. The picture is dubbed to the sound; the sound never regenerates, which is why the lipsync in a four-minute chain never drifts.

One thing we found and fixed: every time a clip is written out through the AAC encoder the trailing partial audio frame is dropped, and the pack then stretches the source audio to the frame-derived length. That is a linear drift of about 1.3 ms per second on the prior soundtrack, up to 32 ms per extension, and a thirty-clip film carries it. Our chain puts the untouched master back over the picture as PCM after every clip and extends that file. Measured: cross-correlation 1.000 in every window.

Job four: scout cheap, finish once. Three previews at 0.5 MP (960x544), three seeds, everything else identical: about four minutes for all three on the 5090. Pick one. Only the pick goes through the 3D latent upscaler to 1.5 MP (1664x928) and a short 8-step er_sde refine at 0.4 denoise: six to eight minutes.

The frame math for a ten-second request, at 24 fps on H3’s 17k+5 grid: 243 frames come out of the model, 39 are the overlap, 204 are new, which is 8.5 seconds of timeline per generation.

The settings

modelminimax_h3_ref2va_pruned_int8_convrot + the 8-step turbo LoRA at 1.0
encoder / VAEsQwen3-VL-32B nvfp4; video VAE fp16, audio VAE fp32
previews960x544, res_multistep / simple, 8 steps, sigma shift 12 / 3, seeds s, s+1, s+2
finish3D latent upscale to 1664x928, er_sde / linear_quadratic, 8 steps at 0.4 denoise
overlap39 frames (1.625 s). The valid ladder is 17n+5: 5, 22, 39, 56
audiolocked master, sliced positionally, remuxed as PCM after every clip
identityour H3 LoRA at 0.7, a head-only face plate, a full-body plate on a grey wall
lighta face-light clause in every scene (“soft key light on her face; exposure set for her face, not the room”)
postnone

The office film, and why it works

The strongest film we made is thirty-one seconds in a night office: a security camera high in a corner, a huge antlered shadow crossing the frosted glass behind her, a low handheld creep toward the partition, and a reverse angle for the reveal (a ginger cat on the photocopier, a pigeon on the sill, one desk lamp). Three clips, three angles, and every angle change lands on a join. Across the whole night’s rendering twenty of twenty-seven seed candidates came back clean; the seven that did not were seed problems on identical text, which is the whole argument for rendering three and picking. Two of the nine clips had exactly one clean seed in three.

It works for the least flattering reason: it has the least going on. One person, one action per clip, a slow camera, and the scares are shadows, which a video model draws beautifully because a shadow has no anatomy.

The direction rules that came out of five films: one agent and one action per clip; the world or the camera angle changes at the join, never mid-clip (mid-clip, the model cuts, and no prompt stops it); direct the end framing rather than the move (a “slow push-in on her face” ends in an extreme close-up, because H3 keeps pushing until the clip ends); and judge a clip-one candidate on its opening, because the first two seconds are the seed’s.

Five ways it broke, and the picture that fixed each one

  1. The car through the door. At the first join of a chase the camera had to get from inside the car to outside, and it went through the door panel. Nothing in the model knows a door is solid. Fix: do not ask for it.
  2. The pursuers that never appeared. Two black SUVs were written into the chase in plain words, headlights blazing, closing in the mirror. They never appeared, in nine candidates. Text alone does not put a second thing in the scene. A picture of the car, given a job in the prompt as <Picture 3>, and the same three clips rendered again with nothing else changed: headlights in the mirror, the car in the tracking shot, three cars through the intersection.
  3. The dance studio that became a home office. Sixteen seconds into a dance, during a floor roll, the studio turned into a desk, a monitor and a lamp. Nobody wrote a home office. The body reference was a waist-up shot at a desk, and the moment her body filled the frame the model fell back to the only room it had been shown. Fix: the body reference rendered again on a plain grey wall, nothing else in frame. The studio holds for all 27 seconds.
  4. The reference played back as a shot. In the second chase, at the last join, all three seeds cut to the reference picture itself, the desk, the monitor, the microphone, for two seconds, inside a car chase. Same cause as 3, nastier form.
  5. The lanyard. The face reference was a crop that included a neckline, so it had a grey T-shirt and a black lanyard in it, and both appeared on her in every chain regardless of what the body picture wore. Fix: crop the face reference at the chin.

Plate rules are identity rules

Four things about the reference pictures decide whether the film looks like the person, and all four were measured.

Brightness. Measured on skin pixels only, the reference face sits at 53 on a 0-100 scale; the same face in the film sits at 40 in a room lit the same way. A lighting clause in the prompt moved it one point. The identity LoRA in the video model moved it three. A colour grade in post closed the gap to three points and read, in our operator’s words, as an obvious filter, so there is no grade. The fix that survived is to write scenes with light on the face and accept that a night scene reads darker than the plate.

Height. In the first dance she read short. Height is decided by the body picture, so we ran a seven-arm ladder on the same seeds. The words tall, 178 cm, long legs, model proportions did nothing, because the identity LoRA’s trained body outvotes every adjective. A 1:2 canvas made her wider. What moved her: the identity LoRA turned down from 1.0 to 0.6 with the words in. A body with no identity LoRA is tall, in her clothes, with a stranger’s face, and even though the prompt says the face picture owns the face, the stranger bleeds into it by six seconds. No donor bodies.

Clothes. In athletic, street or fitted clothes she comes out more accurate, because the identity LoRA was trained on fitted looks and loose fabric buries what it learned to draw. Same seeds, same wall, only the clothes moving: the athletic rows carry the strongest face. Beachwear broke full-body framing three times in four, so it is a face reference, not a body plate.

The crop. Head only, or the clothes come with the face.

So, for a reference set: the face picture is head only; the body picture is on a neutral wall with nothing else in it; fitted clothes, the ones you want on screen; identity strength 0.6 with the proportions written in. None of it is prompt craft. It is all in the pictures.

The identity LoRA, and the ear

Version one of our H3 identity LoRA was trained on 39 photographs and 42 short voice clips. Version three took 456 captioned images already made for an earlier LoRA on a different model, and was better on every axis we could see: no grey T-shirt from the seed, no reference played back as a shot, all three seeds cut-free where v1 cut on two. It was also, when we read the training log properly, images only: the clip encoder rejected the first video clip because its height was not a multiple of 32 and stopped there, before the remaining clips and every voice file; the trainer ran anyway with one placeholder voice step per epoch. Fixed, cache cleared, trained again with everything in it (v3b, 1 h 49 min).

Then the voice. One spoken line rendered with each version; the spectrum said v3b rolled off above 3 kHz where v1 reached past 9 kHz, so we wrote down that 456 photos had diluted 42 voice clips eleven to one. Our operator listened and said v1 and v3b sound roughly the same and v3 is far off. The spectrum was a flag; the ear was the gate. v3b is the studio’s LoRA for picture and voice.

One more thing the LoRA does: with it off, one seed played the head reference back as a full-frame close-up and the face warped in motion. With it on, no seed did. Keep it in.

Instruments over eyes, eyes over instruments

Our first film was reported as cut-free by a detector sampling at 12 fps. Our operator watched it and said it cuts. He was right: two hard cuts, both inside clip two, one a match cut behind a door frame. Between 12 fps samples the ordinary motion of a walking person doubles while a one-frame jump does not, so the jump hides in the noise: the same cut is 10x the median at 24 fps and under 5x at 12. The detector now samples at the native rate and calls a cut a one-frame spike more than 5x the median of the twelve frames either side, more than 20 absolute, and isolated (more than 2.5x both neighbours). It was checked against every known cut we have.

What the detector still cannot see: a sustained morph (the studio becoming an office over 1.5 s) drifts rather than spikes. And a tool we wrote to measure how many heads tall a body plate is printed 13.6 for one of them, which no human body is, because its landmark moved with the hair. Height was judged by eye and we say so.

The bill

Measured on one RTX 5090 at the settings above; your numbers will be yours.

one H3 clip, 5 s~25x its own length to render
10 s~40x
25 s~87x
the chain, 3 previews + 1 finish per clip76-78x real time across four films
a minute of finished filmabout an hour and a quarter, with a three-way choice at every step
the finishgrows with the chain: 380 s on clip 1, 520 s by clip 3 (the loader decodes the whole prior film each time)

The trap: a 0.5 MP preview fits beside another model on the card; the 1.5 MP finish does not. With an image model left resident on a second ComfyUI instance holding 11 GB, the finish did not fail, it paged: one sampling step took 6127 seconds. Empty the card before a finish.

Two things we have not settled, reported rather than ours. The author of the preview workflow says in his own video that a single pass at full resolution beats seed hunting plus the upscaler for final quality and is not slower. We rendered one of our presenter shots both ways at the same seed and put both on screen. Two things to know before you try it: the same seed at a different latent size is a different shot, not the same shot sharper; and on the 5090 the single pass took 764 s for one take against 452 s for the finish after a ~4-minute three-way preview, so on this card it was not cheaper, and it comes with no choice. And he says the valid overlaps run 5, 22, 39 and 56 frames and recommends 22, especially when a voice runs across the join; we use 39 because we copied it. That is the next test.

If your card will not carry this, the same nodes run on Civitai and you pay in Buzz instead of render hours.