Words as Coordinates
For decades, a word was an arbitrary ID. cat and dog were as unrelated as cat
and democracy. Then came a simple, radical reframing:
You shall know a word by the company it keeps.
That is the distributional hypothesis. Words used in similar contexts mean similar things. If you train a model to predict a word’s neighbours, the internal numbers it needs become a map of meaning.
Two methods, one idea
word2vec (2013) offered two training tricks — predict a word from its neighbours, or neighbours from a word. Both produced word embeddings: dense vectors a few hundred numbers long, where cosine similarity means semantic similarity. GloVe arrived a year later with a different counting method and the same destination.
The famous party trick was vector arithmetic. king − man + woman landed almost
exactly on queen. Nobody designed that; it fell out of training. It was the first
clear sign that a network could capture relationships, not just labels.
From words to everything
If a word can be a vector, so can a sentence, a document, an image, a user. The same machinery — train a network, keep the internal representation — turns any object into a point in a shared space. That space is where modern retrieval lives.