Making · 2026
Language models on six gigabytes
Running quantized models on consumer VRAM — working out what Q4_K_M actually costs in quality, and reading enough of nanoGPT to know why attention works.
The constraint
A GTX 1660 Super has 6 GB of VRAM. So does the RTX 4050 in my laptop. That number decides everything about what you can run locally, and reasoning about it properly is most of the skill.
A 7-billion-parameter model at 16-bit precision needs roughly 14 GB for weights alone, before any KV cache. It doesn’t fit, and it isn’t close. At 4-bit it’s around 4 GB, which fits with room for context. Quantization isn’t an optimization here — it’s the difference between running the model and not.
What quantization actually costs
Q4_K_M in the GGUF format is where I’ve landed for almost everything, and understanding why
it works better than it sounds like it should was the interesting part.
Naive 4-bit quantization — one scale factor across a whole tensor — is genuinely destructive,
because a single outlier weight stretches the range and everything else loses resolution. The
k-quant methods quantize in small blocks with per-block scales, so an outlier only degrades its
own neighbourhood. The _M variant then spends extra bits specifically on the tensors that are
most sensitive to precision loss, and fewer on the ones that aren’t.
The result is a model at roughly a quarter the size where the quality difference is, for most work, hard to notice. Where it is noticeable is long-chain reasoning and exact recall — small per-token errors compound over a long generation in a way they don’t over a short one.
Reading the thing rather than using it
Running models locally made me want to know what was actually happening inside them, which led to working through nanoGPT — small enough to hold entirely in your head, complete enough to be the real thing.
The concept that changed how I think about these models is induction heads: attention heads that learn to find an earlier occurrence of the current token and copy what followed it. It’s a simple, almost mechanical circuit, and it’s a large part of what in-context learning is. When a model picks up a pattern from your prompt and continues it, that’s not an abstract emergent mystery — it’s a specific circuit, and you can find it.
That reframing is the most useful thing I’ve gotten out of this. These are not black boxes. They are extremely complicated boxes with findable parts inside.
Setup
LM Studio for quick iteration, llama.cpp directly when I want control over layer offloading
and context. Qwen2.5 and Qwen3-Coder at Q4_K_M for most things, Gemma locally alongside them.