No Free Lunch
A troubling theorem sits at the foundation of learning: averaged over all possible problems, every learning algorithm performs exactly as well as random guessing. The no free lunch theorem.
So why does machine learning work at all? Because the real world is not all possible problems. It has structure — locality, smoothness, repetition — and we build models whose assumptions match it.
Inductive bias: the assumption you cannot avoid
Every model brings an inductive bias — a built-in preference for some patterns over others. A convolutional network assumes nearby pixels matter together. A recurrent network assumes order matters. A transformer assumes anything can attend to anything. Choosing an architecture is choosing which assumptions to bet on.
Occam’s razor is the practical version: among models that fit equally well, prefer the simpler. Simplicity is not an aesthetic — a simpler model has fewer ways to accidentally fit noise, which is why it usually generalizes better.
Capacity is a dial
Capacity is roughly how many patterns a model can fit. Too little and it cannot represent the truth; too much and it can memorize nonsense. Modern deep learning loves huge capacity — and then uses regularization and huge data to keep it honest.