Companion episode — we didn’t dodge the obvious rival; we installed Duix and ran it.
What Duix is
Duix (formerly HeyGem) is the purpose-built open-source answer to HeyGen. It clones a digital human from a short video clip, then drives that clone with any audio you give it — offline, on your own NVIDIA card. If you know this space, it’s the first name people bring up, so we tested it head-to-head.
- Runs offline on an NVIDIA GPU (RTX 4070-class minimum).
- Deploys via Docker + the NVIDIA container toolkit — a container serves a small API; you POST an audio path and a source-video path, it returns a lip-synced clip.
- License: free for commercial use until you’re a big operation (roughly 100k users / $10M revenue) — read the repo’s terms.
What we found
Fed a clean character source, Duix held identity beautifully — a full clip, locked face, clean lip-sync. On the pure “be a specific face and talk” job, it’s genuinely strong. The key insight from testing it:
Duix doesn’t invent motion — it copies the mannerisms of whatever video you feed it, and adds a light polish.
Give it a stiff source and it faithfully copies the stiffness. Give it a lively, gesturing source and it matches the gestures and cleans things up a little. So its ceiling is set by your source clip.
The honest asterisk: we cloned an AI-generated character, not a real person — so it inherited our pipeline’s look, which is not Duix’s true ceiling. Point it at a real human and you’ll get better results than we did.
Where it fits
Here’s the real distinction: Duix re-drives a talking head from a source you already have. It doesn’t generate a body from audio. So it sits in the same slot as a mouth-driver — a finishing tool on top of footage — not as the thing that creates the performance. For a full-body local host, you still want a body generator like LongCat; for a rock-solid talking-head clone of a specific person, Duix is excellent.
Honest verdict
The best-in-class open-source HeyGen clone for identity-locked talking heads, and the honest thing to reach for if your goal is “make this exact face say these exact words.” Just remember what it is: a mirror-and-polish of your source, at its best on a real human — not a body-from-audio generator.
