AI Foundationspredict · compress · act
act XI

The Open-Weight World

The point of open weights is that you can run them. Here is the arithmetic that decides what fits, the tricks that make it fit, and what you give up.

50

Running One Yourself

decadeas of 2026-09
before this →Licences and Rules

Everything so far has been about who builds models and on what terms. This chapter is the payoff: a model on your own machine. The single number that decides whether that is possible is memory, and it follows from one piece of arithmetic.

MODELFULL PRECISION (2 bytes)4-BIT (about 0.5 bytes)3B7B30B70B6 GB14 GB60 GB — data-centre GPU140 GB — several GPUs≈2 GB≈4 GB≈17 GB≈35 GBplus the KV cache, which grows with context length. a long chat can cost more than the weights.
the memory rule. A model needs roughly its parameter count times the bytes per number, plus room for the context you feed it.

The arithmetic

A model stores its weights as numbers. Full precision (16 bits) uses two bytes per number, so a 7-billion-parameter model needs about 14 GB just to sit in memory. Store each number in 4 bits instead and the same model fits in roughly 4 GB. That single decision — how many bits per weight — is what moves a model from a data centre to a laptop.

On top of the weights you pay for the KV cache, which grows with the length of the conversation. It is the reason a model that fits is still slow with a long context, and why the attention trick from DeepSeek exists.

The quantization ladder

Quantization (met earlier in the serving chapter) means storing weights in fewer bits. The usual rungs: 16-bit for training and reference, 8-bit for a near-lossless shrink, and 4-bit for a large shrink with a small, often acceptable loss of quality. Quantized weights are the 4-bit kind you actually download for a home machine. Two formats dominate: GGUF, used by llama.cpp, which runs well on a CPU or a Mac; and AWQ or GPTQ, used for GPU serving.

The software you actually run

pick by what you have

llama.cpp and Ollama are the easy path: run a GGUF file on a laptop, with Ollama handling downloads and a simple API. LM Studio puts a graphical face on the same idea. vLLM and SGLang are for GPUs and throughput — many requests at once — and are what a small service would run. This is local serving: the same job the earlier serving chapter described, done on hardware you own.

a subtlety with mixtures of expertsA sparse model may activate only a few experts per token, but all of them must be in memory. A 671B model needs room for 671B — even if each token only uses 37B. Sparsity saves compute, not space.
DOWNLOADa quantized GGUF or safetensors file
SERVEllama.cpp, Ollama, vLLM, SGLang
ADAPTprompt it, or fine-tune with LoRA

Changing the model, not just the prompt

If prompting is not enough, you can fine-tune. Full fine-tuning updates every weight and needs serious hardware. LoRA freezes the model and trains a small pair of low-rank matrices alongside it — a few percent of the parameters — which is cheap enough to do on one GPU. The rule of thumb: try prompting, then retrieval, and only then fine-tuning. Most problems people reach for a fine-tune to solve are actually context problems.

What you give up

Running locally trades convenience for control. You gain privacy, no per-call cost, no rate limits, and no one changing the model under you. You lose the frontier model’s raw capability, and you take on the operations: updates, monitoring, safety, and the electricity bill. The next and final chapter is about where this whole open world is heading.

introduces →GGUFLoRAlocal servingquantized weights
← previousLicences and Rulesnext →The Open Frontier