Why Memorizing Fails
A model can always score perfectly on data it has already seen — just memorize it. The test is data it has never seen. That gap is where all the interesting failure lives.
The three failures
Underfitting — the model is too simple to capture the pattern. It does badly everywhere. Overfitting — the model has memorized noise as if it were signal. It does beautifully on training data and poorly on new data. The goal is the middle: the simplest model that still explains the data.
The discipline of holding data back
You split your data: a training set the model learns from, and a test set it never sees until the end. The test score is the only honest one. If you tune against the test set, it quietly becomes a second training set and stops telling you the truth; that is why serious work keeps a third, untouched validation split.
Regularization is any pressure against memorization: weight decay, dropout, early stopping, or simply using less capacity. All of them deliberately make the model worse on training data in the hope of making it better on the truth.
The bias-variance tradeoff
Bias is error from a model too rigid to see the pattern. Variance is error from a model so flexible it chases noise. You cannot usually reduce one without raising the other — you balance them.