A gaming laptop running a local AI model with most of the model not on the graphics card

Local AI on 8 GB of VRAM: What Actually Fits, and What Actually Runs

We took a stock Asus TUF gaming laptop with an RTX 3070 and 8 GB of video memory, and benched eight local models overnight across ten tasks. The result that surprised us: a 30 billion parameter model with two thirds of itself off the card beat a 14 billion parameter model that fit better, by nearly four times. File size does not predict speed. All 103 measured rows are free below, failures included.

September 11, 2026 · 6 min · Aillex / DIY AI
A single consumer graphics card driving a video model that produces both picture and sound

MiniMax H3 on One Gaming Card: Five Findings From a Day of Tests

We ran MiniMax H3 for a full day on a single RTX 5090: talking shots driven by our own voice track, a walk and talk, text to speech, a trained voice, and an eight scene short film with its own soundtrack. Every number here came off our machine. So did the mistakes.

September 5, 2026 · 7 min · Aillex / DIY AI
LTX-2.5 in ComfyUI: The Hidden Prompt Rewriter That Poisoned Two Rounds of My Benchmark

LTX-2.5 in ComfyUI: The Hidden Prompt Rewriter That Poisoned Two Rounds of My Benchmark

We benched LTX-2.5 for two full rounds and got confident, wrong answers, because the official templates ship an LLM prompt rewriter that is on by default and buried inside a subgraph. Here is how to find it, how to actually turn it off, and what the model does once your prompts reach it intact.

August 14, 2026 · 8 min · Aillex / DIY AI
As an Amazon Associate, this site earns from qualifying purchases. Referral links are always disclosed.