Memory Outside the Weights
A model’s weights are frozen at training time. It cannot know your documents, your inbox, or yesterday’s news. Retrieval-augmented generation (RAG) fixes this by fetching relevant text at query time and putting it in the context window.
The pipeline
Documents are split into chunks, each chunk is embedded into a vector, and the vectors are stored in a vector database — an index that answers “find the nearest vectors” fast. At query time the question is embedded the same way, the nearest chunks are retrieved, and they are pasted into the prompt. The model then answers from the provided text.
Where it breaks
RAG fails in ways that are easy to miss. Chunking splits ideas mid-sentence, so the retrieved text is missing context. Semantic search finds text that is similar rather than text that answers. The right passage may rank below a merely topical one. And if the answer is not in the corpus, retrieval will confidently return something adjacent, which the model may then present as fact.
Hybrid search — combining embeddings with keyword search — and a reranking step usually fix most of this. So does giving the agent a tool to search, rather than always retrieving before it thinks.
Weights vs. context
Parametric memory is baked into the weights: broad, slow to update, hard to cite. Retrieval memory lives outside: precise, instantly updatable, auditable. You want both — knowledge in the weights, facts in the index.