Update, 30 September 2026. The brain behind Aillex’s live face has moved off the RTX 5090. It now runs Qwen 3.5 9B on a second PC with two 2016 graphics cards, which frees the 5090 for her face. The 26B recipe below is unchanged, and the latency rules apply to any size. How that second PC is set up.
Everything else in a local companion is replaceable; the brain is the soul. This guide runs a 26B-parameter multimodal model served by Ollama, big enough for real conversation and tool use, small enough to leave VRAM for everything else on a 32 GB card. Wondering whether a different model should be your brain? We ran three contenders through the local model gauntlet on real studio jobs, measured.
Sizing honestly
A working local model shelf: each line is a brain you own, sized in gigabytes, no subscription attached.
| Model size (4-bit) | VRAM ballpark | Verdict for a 32 GB card |
|---|---|---|
| 7 to 8B | ~5 to 6 GB | fast, fine for chat, shallow on nuance |
| 13 to 14B | ~9 to 10 GB | the sweet spot for smaller cards |
| ~26B | ~17.6 GB | the pick when the brain shares a 32 GB card with lip-sync + STT |
| 70B | 40 GB+ | does not fit; don’t believe optimistic blog math |
The three latency rules
1. Pin the model. Ollama unloads idle models; a cold load costs many seconds mid-conversation.
curl http://localhost:11434/api/generate \
-d '{"model":"YOUR_MODEL","keep_alive":-1}'
2. Keep the session warm. If your agent layer re-launches a CLI per message, you pay prompt re-ingestion every turn (~8s for a large system prompt). A persistent process holding one session took our brain latency to 0.7 to 3s, the model’s prefix cache does the heavy lifting.
3. Watch reasoning modes. Many modern models default “thinking” on. Great for hard problems; terrible when a quick description burns 23 seconds producing an empty reply. Toggle thinking off for real-time paths.
3b. Sometimes the fix is a purpose-built variant. When we benchmarked our brain against newer models, the thinking-mode version of the challenger was 3 to 7× slower to first token, unusable live, but a no-think variant of the same model matched or beat the incumbent on every check and returned cleaner structured output. In Ollama you can bake this in: create a named model from the base with a system-prompt suffix that disables reasoning, and set num_ctx explicitly at creation, context size is baked at model-create time; a config-file setting alone may load you an 8k-context brain that silently fails tool calls. Verify what actually loaded via /api/ps (look for context_length). On a 32 GB card, 64k context on a ~27B model leaves ~11 GB headroom; 128k would OOM.
Free vision (the multimodal dividend)
If your brain model is multimodal, your assistant can see at zero extra VRAM: screenshot → the already-loaded model describes it → inject the description into conversation context as text. Trigger it on demand (“look at my screen”) rather than continuously, no idle GPU burn. Warm describe on our stack: ~1.2s.
Memory and personality
Raw LLMs forget everything between sessions. The agent layer on top gives Aillex tools and a memory you can edit: her personality and reference files, updated with the points worth keeping (ours also handles MCP tool calls). Two hard-won notes:
- Personality lives in the system prompt, but voice formatting is its own instruction. A chat persona happily emits markdown, emoji and kaomoji, which a TTS then reads aloud. Add an explicit “plain spoken sentences, normal punctuation, no formatting” override for the voice path, and strip residual markup in code.
- Emotion tags are cheap and powerful. We ask the brain to prefix replies with
[happy],[concerned], etc.: one regex later, the avatar has synchronized facial emotion, glow accents and gestures.
One brain, many faces
Point every surface at the same warm brain: our web avatar and a Discord voice bot are thin front-ends to a single session: one memory, one personality, wherever you talk to her.
brain endpoint (one warm session)
├── web avatar page (mic + 3D character)
├── Discord voice bridge
└── anything else that can POST text
Hardware sizing for all of this: the exact PC we run on →
Serving the same kind of brain to a whole household instead: your family’s own ChatGPT on an old gaming PC, one account and one set of instructions per person.
This brain powers everything on YouTube → @AskAillex. Give it ears and a mouth: the full architecture.
