Companion episode — where Eros sits on the honest map of local AI video tools.

What Eros is

Eros (sometimes “10Eros”) is a distilled, baked variant of LTX-2.3 — one local checkpoint that generates synced audio and video from a text prompt, in a single model, in about 8 steps. Feed it a prompt and it invents a scene and a matching soundtrack. It’ll also honor a character image (image-to-video). It’s genuinely impressive that this runs on your own GPU.

The catch nobody mentions

Eros writes its own audio. The soundtrack is generated from your prompt — describe “soft koto music” and you get koto. That means the one thing you probably want — feeding it your own voice and having the character lip-sync to it — it does not do. It’s text-to-audio-video generation, not a dubber. You write what it says and sounds like; you don’t hand it a voice track.

So on our map, Eros lives in the generation column with LTX and Wan — not the talking-avatar column.

Running it (the honest friction)

  • It’s a big baked checkpoint (~46GB). On a 64GB-RAM machine it will exhaust memory and hard-crash mid-render if you’re not careful — free the text-encoder before the diffusion model loads (disable “smart memory”) and keep the frame count modest.
  • Mind the image-conditioning switch. The example workflow ships with the image bypassed and a demo prompt still in it — so it’ll happily ignore your character and render the example scene. Turn the image conditioning on and write a prompt that agrees with your plate, or it does its own thing.

Honest verdict

A fascinating one-model generation toy — local text-to-audio-video is a real capability, and it can hold a character image. But it invents its own audio, so it’s not how you make your presenter say your words. Cool to know exists; for a talking host, you want a real lip-sync tool instead.