AI Foundationspredict · compress · act
act III

The Instruments

The central tension of learning: fit the data you have without fooling yourself about the data you will get.

11

Why Memorizing Fails

before this →Rolling Downhill

A model can always score perfectly on data it has already seen — just memorize it. The test is data it has never seen. That gap is where all the interesting failure lives.

model complexity →errortraining errortest errorsweet spotunderfitoverfit
the U-shaped truth: too simple underfits, too complex overfits

The three failures

Underfitting — the model is too simple to capture the pattern. It does badly everywhere. Overfitting — the model has memorized noise as if it were signal. It does beautifully on training data and poorly on new data. The goal is the middle: the simplest model that still explains the data.

The discipline of holding data back

You split your data: a training set the model learns from, and a test set it never sees until the end. The test score is the only honest one. If you tune against the test set, it quietly becomes a second training set and stops telling you the truth; that is why serious work keeps a third, untouched validation split.

Regularization is any pressure against memorization: weight decay, dropout, early stopping, or simply using less capacity. All of them deliberately make the model worse on training data in the hope of making it better on the truth.

The bias-variance tradeoff

two ways to be wrong, forever

Bias is error from a model too rigid to see the pattern. Variance is error from a model so flexible it chases noise. You cannot usually reduce one without raising the other — you balance them.

the ruleJudge a model by how it handles the unseen. Everything else is a proxy.
UNDERFIThigh bias — too simple
JUST RIGHTcaptures signal, ignores noise
OVERFIThigh variance — memorized the noise
introduces →overfittingunderfittingregularizationtraining settest setbias-variance tradeoff
← previousRolling Downhillnext →No Free Lunch