This guide is part of the companion episode — the full, honest comparison of local AI video tools.

What LTX-2.3 is

LTX-2.3 (Lightricks) is a local image-to-video and text-to-video model. You give it a starting image and a prompt, and it animates it — 30fps, native portrait and high resolution, and it renders faster than real time on a modern card. It runs entirely on your own GPU inside ComfyUI. No cloud, no queue, no per-second billing.

It’s the model this whole channel started on, and it’s still the one we reach for first when we need motion from an image.

Why it’s our generation workhorse

  • Speed. Faster-than-realtime generation means you iterate in seconds, not minutes.
  • Identity truth. LTX stays truer to your starting plate than most local models — the face you feed in is the face that comes out, which matters enormously for a recurring character.
  • Single-pass coherence. A single-stage LTX gen holds together across a long clip with no identity drift or quality falloff — you don’t have to chain short segments.
  • It’s free and it’s yours. Every frame runs on hardware you own.

The catches (honest)

  • Two-stage adds crunch. The spatial-upscaler two-stage path sharpens too hard — over-baked skin, tendons. Single-stage is softer and is the better base if you’re going to lip-sync on top of it. Upscale the final output later (with a temporally-aware upscaler) instead.
  • Keep motion minimal for a clean base. Big head moves make LTX briefly exaggerate anatomy (neck tendons on a tilt-back). For a lip-sync canvas, blinks-and-breathing beats dramatic motion.
  • Audio-sync is not its job. LTX generates video; getting a mouth to track a specific voice consistently is where it slips, and long good renders take time. That’s exactly why, on our channel, generation and talking became two separate tools.

Where it fits

LTX-2.3 is the generation layer: it makes the moving footage. For a talking presenter, we generate the body/motion with a specialist and let a mouth specialist handle the lips — but the raw, on-model motion underneath is LTX. If you only take one local video model, this is the one.

Honest verdict

The local generation king for 2026: fast, faithful, free. Not a one-click talking-avatar button — pair it with a lip-sync tool for that — but as the engine that turns a still into believable motion on your own machine, nothing else we’ve tested is this practical.