Update, August 2026: this model is one stage of the four-model chain that renders our presenter. The full pipeline is now public: Free Local HeyGen: the four-model chain.
Companion episode: LongCat is the talking-avatar method we actually build on.
What LongCat is
Most lip-sync tools give you a talking head. LongCat gives you a talking body: it drives gestures, lean, head movement, face and mouth from a single audio track. It’s audio-in, animated-character-out, and it runs on your own GPU for free. For a presenter who should look alive rather than pasted-in, that’s a different league.
Why it’s our pick for a full-body local host
- It gestures with the speech. The hands and posture move in tune with what’s being said, the thing that makes a talking clip read as a person.
- Long takes in one job. LongCat uses native sliding windows, so you can request a long clip and it blends the joins internally, we’ve pushed single gens to 40 to 50 seconds without stitching.
- Local and free. No cloud avatar subscription, no per-minute billing.
On the axis that matters for a full-body, expressive character host, we believe this is the best local competitor to HeyGen: local, free, full-body, which nobody else is really shipping.
The honest caveats
- Quiet background helps. Busy or figure-filled backgrounds confuse it, longCat will try to animate anything that looks like a person, including someone on a monitor behind you. Keep the set clean.
- Drift on very long takes. Past roughly 20 to 30 seconds the realism can slowly regress, she gets a touch less photoreal the longer the single take runs. The practical fix is to keep hero takes under ~30 seconds and split longer ones. (This is a real, measurable limit, not a hype-thread “it’s flawless.”)
- The default mouth is too expressive. Out of the box, LongCat’s mouth and jaw over-articulate, cartoonish. We tune ours to calm that down; we’ll show exactly how once the channel hits some subscriber milestones. Until then: know that the raw default mouth is the weak point, and that it’s fixable.
Where it fits
LongCat is the body and performance layer. Feed it a clean character plate and a voice track; get back a gesturing, speaking avatar. It’s the foundation of a serious local presenter.
Honest verdict
The strongest local, full-body, free talking avatar we’ve tested, genuinely a local answer to the paid cloud tools, if you give it a quiet set, keep takes reasonable, and calm the default mouth. That combination is why it’s the one we build on.
Curious about the purpose-built HeyGen clone instead? See Duix / HeyGem, a different tool for a different job.
