The companion episode — the full rig guide, with every board animated.

Almost everyone shopping for an AI computer asks the same first question: how fast is it? Wrong question. The number that decides whether you can run something at all — before speed ever enters the conversation — is memory. Video memory. VRAM.

This guide is the one-page version of the episode: what actually matters, what to buy at every budget in a market that has frankly lost its mind, and when you shouldn’t buy at all.

Disclosure: product links below are affiliate links (Amazon Associate / referral) — they cost you nothing and help keep every guide free. We only recommend what we’d buy ourselves. Full disclosure.

The one number that matters

A model is billions of numbers, and all of them have to live in your graphics card’s memory while it runs. Fits in VRAM → it flies. Spills into system RAM → everything crawls. So think of VRAM as the size of your AI’s desk: a bigger desk doesn’t make you faster — it makes bigger jobs possible.

Two jobs happen when an AI answers you: reading your prompt (raw compute) and writing the answer word-by-word (memory bandwidth — this is why two cards with identical “power” can feel completely different).

The rule of thumb: at everyday (Q4) quality, a model needs a little over half a gigabyte of VRAM per billion parameters.

Model sizeVRAM neededRuns on
8B~5 GBalmost any modern card
32B (the local sweet spot)~20 GBused 3090 / 4090 / 5090
70B~40 GB+dual GPUs, Mac 128GB, or an AI appliance

The 2026 memory crisis (read this before buying anything)

The AI datacenter buildout has drained the world’s memory-chip supply, and GPU street prices have gone haywire. The RTX 5090 lists at $1,999 — on the street it’s routinely $3,000–$4,300+, when you can find one. People call them paper cards: the price exists, the card doesn’t.

In this market the smart moves are: buy the value (the cards still near MSRP), buy used (24GB of yesterday’s flagship), or don’t buy — rent (more below).

The GPU board — sorted by memory

CardVRAMStreet (Jul 2026)The take
RTX 509032 GB~$3,000–4,300the king — and the crisis poster child
RTX 508016 GB~$1,500caught in the price mess
RTX 5070 Ti16 GB~$919–1,200the sane Blackwell 16GB
RX 9070 XT16 GB~$650 (≈MSRP!)the value pick — ROCm 7.2 runs natively on Windows, out-of-the-box Ollama/LM Studio/ComfyUI
RTX 3090 (used/renewed) ★24 GBused marketthe smart money — max VRAM per dollar
RTX 4060 Ti 16GB16 GB~$500squietly beats faster 8GB cards
Arc B58012 GB~$303the entry ticket

Apple silicon is the other road: unified memory means a Mac can hold 128–256 GB — models no single gaming card can touch — but it generates at a comfortable reading pace, not a blur. Buy a Mac to hold one huge model quietly; buy NVIDIA/AMD to iterate fast.

The AI appliance — the new category

Instead of building a PC around a GPU, you can now buy a finished AI box: NVIDIA DGX Spark ($3,999, 4TB) or its cheaper twin ASUS Ascent GX10 ($2,999, 1TB) — the same GB10 Grace Blackwell chip, 128 GB of unified memory in something you hold in one hand.

The honest trade: that memory is slower than a GPU’s, so it generates at reading pace. You don’t buy one for fast video — you buy it to hold giant 70–120B models and run always-on agent teams, quietly, sipping power. You can even stack two. (It’s on our own wishlist.)

Match the rig to the dream

You want to…NeedsThat means
Chat, write, code8–12 GBany modern 12GB card
The genuinely smart 32B models~20 GBused 3090, or 5090 headroom
Make images (SDXL → FLUX)12–24 GB12GB entry, 24GB for the best
Make video 🐉16–24+ GBthe hungriest thing on this channel

Chat is light. Pictures are medium. Video is the dragon. Buy for the heaviest thing you’ll actually do — not the heaviest thing that exists.

The four builds

BuildGPURAM~TotalRuns
Starter16GB (4060 Ti / 9070 XT)32 GB~$900–1,000mid models, images, audio
Creatormore VRAM64 GB~$1,500smooth images, entry video
Pro24GB (used 3090 / 4090)64–96 GB~$2,000–3,00070B models, full-quality video
Dreamdual high-end128 GB$4,000+agent fleets — you don’t need this

Most people should stop at Starter — and that’s correct. Prefer a prebuilt? Our own rig is an iBuyPower — links on the start-here guide.

When the model doesn’t fit

  1. Make it fit anyway — offloading keeps part of the model on the GPU and parks the rest in system RAM. Slower, but it runs. This is why a 64GB RAM kit is the best quiet upgrade in local AI.
  2. Rent itRunPod (referral — new accounts get a one-time credit bonus) spins up a real RTX 5090 for ~$0.99/hr, billed by the second. Perfect for trying before buying, weak PCs, or keeping your GPU free for gaming.

Buy vs rent — the honest math

Renting a 5090 ≈ $1/hr. Buying one at street price ≈ $3,000. That’s roughly 3,000 hours (~126 days of nonstop use) to break even — before electricity.

  • Buy if you’ll run it daily for months and you value privacy and never seeing a meter. It’s yours forever — that’s the whole ethos here.
  • Rent if you’re starting out, unsure, on a weak PC, or the card you want is literally out of stock. Right now, rent-first is genuinely smart money.

Prices verified July 2026 — this page gets updated as the market settles. Start building at askaillex.com, and bring questions to r/aillex.