AI Foundationspredict · compress · act
act IX

The Agent Infrastructure

Training is a one-time cost. Serving is forever. Inference engineering is where the money and the latency live.

38

How Tokens Get Served

before this →Many Hands, One Job

A trained model is a giant pile of matrix multiplications that must run every time someone types a character. Inference is that process, and making it fast and cheap is a field of its own.

“The”“cat”“sat”“on”keyvaluecached — never recomputedeach new token:one query, all keyswithout the cache, cost growsquadratically with length
the KV cache: store past keys and values so each new token is one step, not N

The bottleneck nobody expects

Generating text is memory-bound, not compute-bound. To produce one token, the model must read its entire weight matrix — hundreds of gigabytes — and use each number exactly once. The GPU sits idle waiting for memory. This is why batching matters so much: running many requests together amortises the weight read across all of them, and throughput scales almost for free.

The KV cache is the other half. Attention needs the keys and values of every previous token. Recomputing them each step would be quadratic, so they are stored and reused. The cost is memory: long conversations need huge caches.

The toolkit

vLLM and similar servers introduced paged attention, which manages the KV cache efficiently and enables continuous batching — new requests join mid-flight instead of waiting for a batch to finish. Quantization shrinks weights from 16 bits to 8 or 4, cutting memory and speeding up reads at a small accuracy cost. Speculative decoding uses a small draft model to guess several tokens ahead; the big model verifies them in one pass, turning serial generation into parallel work.

Latency vs. throughput

two different products

Throughput (tokens per second across all users) is a batched, data-centre concern. Latency (time to first token for one user) is a user-experience concern. Techniques that help one can hurt the other — bigger batches improve throughput and worsen latency.

the economicsServing cost is dominated by memory and KV cache, not raw FLOPs. Capability per dollar is an architectural choice as much as a model choice.
PREFILLread the prompt — compute-heavy
CACHEstore keys and values once
DECODEemit tokens — memory-bound, batch it
introduces →inferenceKV cachebatchingquantizationvLLMspeculative decoding
← previousMany Hands, One Jobnext →A Safe Place to Act