How Tokens Get Served
A trained model is a giant pile of matrix multiplications that must run every time someone types a character. Inference is that process, and making it fast and cheap is a field of its own.
The bottleneck nobody expects
Generating text is memory-bound, not compute-bound. To produce one token, the model must read its entire weight matrix — hundreds of gigabytes — and use each number exactly once. The GPU sits idle waiting for memory. This is why batching matters so much: running many requests together amortises the weight read across all of them, and throughput scales almost for free.
The KV cache is the other half. Attention needs the keys and values of every previous token. Recomputing them each step would be quadratic, so they are stored and reused. The cost is memory: long conversations need huge caches.
The toolkit
vLLM and similar servers introduced paged attention, which manages the KV cache efficiently and enables continuous batching — new requests join mid-flight instead of waiting for a batch to finish. Quantization shrinks weights from 16 bits to 8 or 4, cutting memory and speeding up reads at a small accuracy cost. Speculative decoding uses a small draft model to guess several tokens ahead; the big model verifies them in one pass, turning serial generation into parallel work.
Latency vs. throughput
Throughput (tokens per second across all users) is a batched, data-centre concern. Latency (time to first token for one user) is a user-experience concern. Techniques that help one can hurt the other — bigger batches improve throughput and worsen latency.