A gaming laptop running a local AI model with most of the model not on the graphics card

Local AI on 8 GB of VRAM: What Actually Fits, and What Actually Runs

We took a stock Asus TUF gaming laptop with an RTX 3070 and 8 GB of video memory, and benched eight local models overnight across ten tasks. The result that surprised us: a 30 billion parameter model with two thirds of itself off the card beat a 14 billion parameter model that fit better, by nearly four times. File size does not predict speed. All 103 measured rows are free below, failures included.

September 11, 2026 · 6 min · Aillex / DIY AI
The Local Model Gauntlet: gemma4 vs qwen3.8 vs muse-glimmer

The Local Model Gauntlet: gemma4 vs qwen3.8 vs muse-glimmer

Three open-weight models, one RTX 5090, and the same jobs: a real bug, playable games, tool calling, long documents. Measured numbers, two defaults that were quietly breaking the results, and the finding we published wrong and had to correct.

August 24, 2026 · 11 min · Aillex / DIY AI
As an Amazon Associate, this site earns from qualifying purchases. Referral links are always disclosed.