This is HeyGen-style talking-head video from one static photo, free, on your own PC. Every talking shot on this channel is four AI models in a row. We have never said which ones, or in what order, and viewers have been watching the result work for over a month without being told how.
This is the whole thing. The models, the order, and what each one is there to fix.
We are holding one category back: the tuning values. Strengths, thresholds, the blend numbers we arrived at by burning a lot of GPU hours. Those stay ours. Everything structural is here, and the structure is the part that teaches, because the insight is not any single model. It is the order.
Why a chain instead of a model
Every local talking head tool does one job well and leaves a different mess behind.
Drive a still image with audio and you get accurate lip movement attached to a body that sits there like a photograph. Use a video model instead and you get lifelike motion with a mouth that is roughly, vaguely, not really saying your words. Pick either and you are choosing which flaw your viewers notice.
The chain exists because each tool’s leftovers happen to be the next tool’s specialty.
The chain
Stage one: LongCat. Starts from a still plate of the character and produces the take. Body motion, gesture, head movement, the small shifts that make someone look alive rather than pasted in. The mouth at this point is approximate. It moves, but it is not saying your lines.
Stage two: MuseTalk. Repaints the mouth region to the actual recorded voiceover. This is where the words arrive. What it leaves behind is a jaw that overstates itself, especially on wide vowels, which reads as slightly rubbery once you have watched it enough times to know what to look for.
Stage three: Duix. A talking head model in its own right, used here as a cleanup pass over the face rather than as the engine. It settles what stage two exaggerated.
Stage four: MuseTalk again. The second pass is not a repeat. Running it over the cleaned face recovers the precise mouth shapes, the visemes, that the polish pass smooths away. Jaw exaggeration goes in, correct viseme detail comes out.
Nothing here is exotic. All four are open and run locally, all four have their own guides on this site, and the arrangement is the only original thing about it. That is genuinely the lesson: LongCat, MuseTalk and Duix were each documented here months ago as separate tools. The chain is what we did with them.
The failures that produced it
We did not design this. We arrived at it by shipping things that were not good enough.
The looped cam. Early episodes reused a short clip of the presenter on a loop under the voiceover. Our own review of it was that it was painful to watch, and that was the polite version. Loops read as fake faster than almost any other shortcut, because human beings are very good at spotting a repeating motion.
The cache that lied. At one point we re-rendered a segment to the same filename and the pipeline quietly served a face from an earlier build, cached upstream. The audio was new, the wardrobe was new, the face was three weeks old, and nothing anywhere reported a problem. If you build a multi stage pipeline, assume any stage may be caching, and verify by something other than the thing you are looking at. We now check the dress and the background, not the face.
Plate drift. Long takes would slowly stop looking like our character. It took us roughly twenty episodes of guessing before we identified the actual trigger, which is that if the character’s head or hair crosses the edge of the frame in the source plate, the take drifts. Keep the whole head and hair inside the frame with clear space above, and it holds.
Why we kept it quiet
A two month old channel does not have much that is genuinely its own. It has no audience, no back catalogue, no brand. What we had was a chain nobody else had bothered to assemble, and for a while that was the only real asset here.
So we set a number, told people there was something we were not explaining yet, and published it when the number arrived. That is more honest than pretending the silence was modesty, and it is a better deal than the version of this you usually see, where the technique goes behind a course.
The tuning values stay private, and we would rather say so than pretend the recipe here is complete. Someone can rebuild the architecture from this page. They will have to do their own tuning, which is fair, because we did ours.
What it costs
Four passes over the same footage, one after another. This is not a real time system and it is not close to one. It is worth it for a presenter you will reuse across dozens of episodes, and it is absolutely not worth it for a single clip. If you need one talking shot this afternoon, use one tool and accept its particular flaw.
Our companion, which has to answer while you wait, uses a completely different and much lighter approach. That is the companion guide, and the trade is exactly what you would expect: it responds in seconds and looks less finished.
Would a newer model replace it
We tested that question last week rather than assuming the answer. The newest open video model with native audio went through a two round bench in an isolated install, and for speech it lost to this chain. The write up of that, including a default setting that poisoned two rounds of our own results, is here.
That is why the chain is still worth explaining rather than replacing. It will be replaced eventually, probably by a single model that does all four jobs at once, and when that happens we will say so on camera the same week we find out.
The full story, with the ugly intermediate renders from every stage, is on the channel. Questions and builds at r/aillex.