The companion episode: the full rig guide, with every board animated.
Almost everyone shopping for an AI computer asks the same first question: how fast is it? Wrong question. The number that decides whether you can run something at all, before speed ever enters the conversation, is memory. Video memory. VRAM.
This guide is the one-page version of the episode: what actually matters, what to buy at every budget in a market that has frankly lost its mind, and when you shouldn’t buy at all.
Disclosure: product links below are affiliate links (Amazon Associate / referral), they cost you nothing and help keep every guide free. We only recommend what we’d buy ourselves. Full disclosure.
The one number that matters
A model is billions of numbers, and all of them have to live in your graphics card’s memory while it runs. Fits in VRAM → it flies. Spills into system RAM → everything crawls. So think of VRAM as the size of your AI’s desk: a bigger desk doesn’t make you faster, it makes bigger jobs possible.
Two jobs happen when an AI answers you: reading your prompt (raw compute) and writing the answer word-by-word (memory bandwidth, this is why two cards with identical “power” can feel completely different).
The rule of thumb: at everyday (Q4) quality, a model needs a little over half a gigabyte of VRAM per billion parameters.
| Model size | VRAM needed | Runs on |
|---|---|---|
| 8B | ~5 GB | almost any modern card |
| 32B (the local sweet spot) | ~20 GB | used 3090 / 4090 / 5090 |
| 70B | ~40 GB+ | dual GPUs, Mac 128GB, or an AI appliance |
The 2026 memory crisis (read this before buying anything)
The AI datacenter buildout has drained the world’s memory-chip supply, and GPU street prices have gone haywire. The RTX 5090 lists at $1,999: on the street it’s routinely $3,000 to $4,300+, when you can find one. People call them paper cards: the price exists, the card doesn’t.
In this market the smart moves are: buy the value (the cards still near MSRP), buy used (24GB of yesterday’s flagship), or don’t buy, rent (more below).
The GPU board, sorted by memory
| Card | VRAM | Street (Jul 2026) | The take |
|---|---|---|---|
| RTX 5090 | 32 GB | ~$3,000 to 4,300 | the king, and the crisis poster child |
| RTX 5080 | 16 GB | ~$1,500 | caught in the price mess |
| RTX 5070 Ti | 16 GB | ~$919 to 1,200 | the sane Blackwell 16GB |
| RX 9070 XT ★ | 16 GB | ~$650 (≈MSRP!) | the value pick, rOCm 7.2 runs natively on Windows, out-of-the-box Ollama/LM Studio/ComfyUI |
| RTX 3090 (used/renewed) ★ | 24 GB | used market | the smart money, max VRAM per dollar |
| RTX 4060 Ti 16GB | 16 GB | ~$500s | quietly beats faster 8GB cards |
| Arc B580 | 12 GB | ~$303 | the entry ticket |
Apple silicon is the other road: unified memory means a Mac can hold 128 to 256 GB, models no single gaming card can touch, but it generates at a comfortable reading pace, not a blur. Buy a Mac to hold one huge model quietly; buy NVIDIA/AMD to iterate fast.
The AI appliance, the new category
Instead of building a PC around a GPU, you can now buy a finished AI box: NVIDIA DGX Spark ($3,999, 4TB) or its cheaper twin ASUS Ascent GX10 ($2,999, 1TB), the same GB10 Grace Blackwell chip, 128 GB of unified memory in something you hold in one hand.
The honest trade: that memory is slower than a GPU’s, so it generates at reading pace. You don’t buy one for fast video, you buy it to hold giant 70 to 120B models and run always-on agent teams, quietly, sipping power. You can even stack two. (It’s on our own wishlist.)
Match the rig to the dream
| You want to… | Needs | That means |
|---|---|---|
| Chat, write, code | 8 to 12 GB | any modern 12GB card |
| The genuinely smart 32B models | ~20 GB | used 3090, or 5090 headroom |
| Make images (SDXL → FLUX) | 12 to 24 GB | 12GB entry, 24GB for the best |
| Make video 🐉 | 16 to 24+ GB | the hungriest thing on this channel |
Chat is light. Pictures are medium. Video is the dragon. Buy for the heaviest thing you’ll actually do, not the heaviest thing that exists.
The four builds
| Build | GPU | RAM | ~Total | Runs |
|---|---|---|---|---|
| Starter | 16GB (4060 Ti / 9070 XT) | 32 GB | ~$900 to 1,000 | mid models, images, audio |
| Creator | more VRAM | 64 GB | ~$1,500 | smooth images, entry video |
| Pro | 24GB (used 3090 / 4090) | 64 to 96 GB | ~$2,000 to 3,000 | 70B models, full-quality video |
| Dream | dual high-end | 128 GB | $4,000+ | agent fleets, you don’t need this |
Most people should stop at Starter, and that’s correct. Prefer a prebuilt? Our own rig is an iBuyPower, links on the start-here guide.
When the model doesn’t fit
- Make it fit anyway, offloading keeps part of the model on the GPU and parks the rest in system RAM. Slower, but it runs. This is why a 64GB RAM kit is the best quiet upgrade in local AI.
- Rent it: RunPod (referral, new accounts get a one-time credit bonus) spins up a real RTX 5090 for ~$0.99/hr, billed by the second. Perfect for trying before buying, weak PCs, or keeping your GPU free for gaming.
Buy vs rent, the honest math
Renting a 5090 ≈ $1/hr. Buying one at street price ≈ $3,000. That’s roughly 3,000 hours (~126 days of nonstop use) to break even, before electricity.
- Buy if you’ll run it daily for months and you value privacy and never seeing a meter. It’s yours forever, that’s the whole ethos here.
- Rent if you’re starting out, unsure, on a weak PC, or the card you want is literally out of stock. Right now, rent-first is genuinely smart money.
Once you have the card, the next question is what to run on it. The Local Model Gauntlet has the measured VRAM and tokens-per-second for three open-weight models on a 32 GB card, plus the row for 16 GB.
Already own an 8 GB card and not replacing it? That is the most common situation and it has its own page: local AI on 8 GB of VRAM, benched on an RTX 3070 laptop. The headline is that a 30 billion parameter model with two thirds of itself off the card beat a 14 billion parameter model that fit better, by nearly four times. All 103 measured rows are free to download there.
Prices verified July 2026, this page gets updated as the market settles. Start building at askaillex.com, and bring questions to r/aillex.
Related
- 86 failures in one published episode — every class of mistake we found making local AI video, and the checks that now catch them.
