Running One Yourself
Everything so far has been about who builds models and on what terms. This chapter is the payoff: a model on your own machine. The single number that decides whether that is possible is memory, and it follows from one piece of arithmetic.
The arithmetic
A model stores its weights as numbers. Full precision (16 bits) uses two bytes per number, so a 7-billion-parameter model needs about 14 GB just to sit in memory. Store each number in 4 bits instead and the same model fits in roughly 4 GB. That single decision — how many bits per weight — is what moves a model from a data centre to a laptop.
On top of the weights you pay for the KV cache, which grows with the length of the conversation. It is the reason a model that fits is still slow with a long context, and why the attention trick from DeepSeek exists.
The quantization ladder
Quantization (met earlier in the serving chapter) means storing weights in fewer
bits. The usual rungs: 16-bit for training and reference, 8-bit for a near-lossless
shrink, and 4-bit for a large shrink with a small, often acceptable loss of quality.
Quantized weights are the 4-bit kind you actually download for a home machine. Two
formats dominate: GGUF, used by llama.cpp, which runs well on a CPU or a Mac; and
AWQ or GPTQ, used for GPU serving.
The software you actually run
llama.cpp and Ollama are the easy path: run a GGUF file on a laptop, with Ollama handling downloads and a simple API. LM Studio puts a graphical face on the same idea. vLLM and SGLang are for GPUs and throughput — many requests at once — and are what a small service would run. This is local serving: the same job the earlier serving chapter described, done on hardware you own.
Changing the model, not just the prompt
If prompting is not enough, you can fine-tune. Full fine-tuning updates every weight and needs serious hardware. LoRA freezes the model and trains a small pair of low-rank matrices alongside it — a few percent of the parameters — which is cheap enough to do on one GPU. The rule of thumb: try prompting, then retrieval, and only then fine-tuning. Most problems people reach for a fine-tune to solve are actually context problems.
What you give up
Running locally trades convenience for control. You gain privacy, no per-call cost, no rate limits, and no one changing the model under you. You lose the frontier model’s raw capability, and you take on the operations: updates, monitoring, safety, and the electricity bill. The next and final chapter is about where this whole open world is heading.