Companion episode — Sonic is the first talking-avatar method we show, and the first we rejected.
What Sonic is
Sonic takes a still image plus an audio track and produces a talking head — the lips move in time with the voice. On paper, exactly what you want. We tried it early, and we kept it in the comparison specifically as the cautionary example.
Why we rejected it
Sonic regenerates the entire face on every frame to drive the mouth. Because nothing is held fixed, everything drifts:
- The eyes roll and go white.
- The irises drift and change color frame to frame.
- The teeth get reinvented every time the mouth opens.
The lips do sync — but the price is a face that quietly falls apart the longer it talks. Tuning the settings softens it; it never fixes it, because it’s architectural: a method that re-draws the whole face can’t leave the eyes alone.
The lesson
For a recurring character, consistency is the entire game. A method that mangles the eyes for the sake of the mouth is unshippable — your audience notices, even if they can’t name what’s wrong. The fix isn’t a better setting; it’s a better idea: only touch the part of the face that needs to move.
Where it fits
Nowhere, for us — but it’s the perfect illustration of the core principle that shapes the rest of the map: repaint the mouth, leave the eyes. That’s what MuseTalk does, and it’s why the good methods look calm while Sonic looks haunted.
Honest verdict
Rejected. Watch the eyes in any lip-sync demo — that’s the tell. If they roll or recolor, the method is regenerating the whole face, and you’ll regret it on a channel.
