A gaming laptop running a local AI model with most of the model not on the graphics card

Local AI on 8 GB of VRAM: What Actually Fits, and What Actually Runs

We took a stock Asus TUF gaming laptop with an RTX 3070 and 8 GB of video memory, and benched eight local models overnight across ten tasks. The result that surprised us: a 30 billion parameter model with two thirds of itself off the card beat a 14 billion parameter model that fit better, by nearly four times. File size does not predict speed. All 103 measured rows are free below, failures included.

September 11, 2026 · 6 min · Aillex / DIY AI
AI Models in Artificial Intelligence, Explained (for People Building at Home)

AI Models in Artificial Intelligence, Explained (for People Building at Home)

What an AI model actually is, what 7B/70B and quantization mean for your GPU, open vs closed weights, the model types you’ll actually use, and how to pick one without a computer science degree.

July 14, 2026 · 3 min · Aillex / DIY AI
Run a 26B AI Brain Locally, Warm, Multimodal, and With Memory

Run a 26B AI Brain Locally, Warm, Multimodal, and With Memory

The brain is the biggest VRAM line-item and the biggest latency trap. How we run a 26B multimodal LLM via Ollama with sub-second warm responses, persistent memory, and free screen vision.

July 1, 2026 · 4 min · Aillex / DIY AI
As an Amazon Associate, this site earns from qualifying purchases. Referral links are always disclosed.