Making · 2026

Language models on six gigabytes

Running quantized models on consumer VRAM — working out what Q4_K_M actually costs in quality, and reading enough of nanoGPT to know why attention works.

llama.cpp · Quantization · Transformers

The constraint

A GTX 1660 Super has 6 GB of VRAM. So does the RTX 4050 in my laptop. That number decides everything about what you can run locally, and reasoning about it properly is most of the skill.

A 7-billion-parameter model at 16-bit precision needs roughly 14 GB for weights alone, before any KV cache. It doesn’t fit, and it isn’t close. At 4-bit it’s around 4 GB, which fits with room for context. Quantization isn’t an optimization here — it’s the difference between running the model and not.

What quantization actually costs

Q4_K_M in the GGUF format is where I’ve landed for almost everything, and understanding why it works better than it sounds like it should was the interesting part.

Naive 4-bit quantization — one scale factor across a whole tensor — is genuinely destructive, because a single outlier weight stretches the range and everything else loses resolution. The k-quant methods quantize in small blocks with per-block scales, so an outlier only degrades its own neighbourhood. The _M variant then spends extra bits specifically on the tensors that are most sensitive to precision loss, and fewer on the ones that aren’t.

The result is a model at roughly a quarter the size where the quality difference is, for most work, hard to notice. Where it is noticeable is long-chain reasoning and exact recall — small per-token errors compound over a long generation in a way they don’t over a short one.

Reading the thing rather than using it

Running models locally made me want to know what was actually happening inside them, which led to working through nanoGPT — small enough to hold entirely in your head, complete enough to be the real thing.

The concept that changed how I think about these models is induction heads: attention heads that learn to find an earlier occurrence of the current token and copy what followed it. It’s a simple, almost mechanical circuit, and it’s a large part of what in-context learning is. When a model picks up a pattern from your prompt and continues it, that’s not an abstract emergent mystery — it’s a specific circuit, and you can find it.

That reframing is the most useful thing I’ve gotten out of this. These are not black boxes. They are extremely complicated boxes with findable parts inside.

Setup

LM Studio for quick iteration, llama.cpp directly when I want control over layer offloading and context. Qwen2.5 and Qwen3-Coder at Q4_K_M for most things, Gemma locally alongside them.

← All making

Contact

Let's build something.

I'm looking for data science, machine learning and AI engineering roles in Phoenix or remote — and I'm always happy to talk about models, retrieval, or why yours is overfitting.