Seeing with Windows
An image has millions of pixels. A plain network would treat every pixel as unrelated to every other. A convolutional neural network (CNN) assumes something obvious and true: what matters is local. It slides a small window across the image, looking for the same pattern everywhere.
Weight sharing
The same small filter is used at every position. That is the key trick: instead of learning a separate weight for every pixel, the network learns a handful of reusable detectors. This is a strong inductive bias — and for images, it is exactly right. The network needs far less data to learn “edge,” “corner,” and eventually “eye” and “wheel,” because those features are translation-invariant.
Early layers find edges. Deeper layers compose edges into textures, textures into parts, parts into objects. Pooling layers shrink the map between stages, so later layers see a wider slice of the image with the same filter size.
The 2012 breakthrough
AlexNet — a deep CNN trained on GPUs — crushed the ImageNet competition in 2012 by a margin nobody expected. It combined three things this guide has already covered: convolution, backpropagation, and GPU parallelism. The modern era starts here.