AI Foundationspredict · compress · act
act VI

The Connectionist Turn

Not all learning has a correct answer. Sometimes there is only a score — and you must figure out which action earned it.

23

Learning from Reward

before this →Words as Coordinates
1992TD-Gammon — a network learns backgammon by playing itself, given no human strategy
2015DQN — One network learns dozens of Atari games from the pixels alone.

Supervised learning needs labels. But how do you label the right move in a game you are still learning to play? Reinforcement learning (RL) works from reward alone, trial and error, over time. It is how agents learn to act.

policy πenvironmentaction astate s′ + reward rmaximise total reward over time
the agent explores, the environment rewards, the policy improves

The core objects

A policy is the agent’s rule for choosing actions — the thing being learned. A value function estimates how good a state is, or how good an action is in that state. The reward signal might be sparse: a single +1 at the end of a long game, with nothing to say which of the thousand moves was responsible. That is the credit-assignment problem again, now spread across time.

Exploration is the central tension. Take the action you think is best, and you never learn whether something better exists. Try new things, and you waste reward. Every RL algorithm is, in part, a policy for balancing the two.

The two great families

value-based and policy-based

Q-learning learns the value of every action and picks the best. Policy gradient methods skip values and push probabilities directly toward actions that led to reward. Modern systems usually blend the two.

the honest caveatRL is powerful and fragile. Reward is easy to specify and hard to specify correctly. An agent optimises exactly what you measured, not what you meant — the flaw we will meet again as reward hacking.
ACTsample an action from the policy
SCOREenvironment returns reward
UPDATEmake rewarded actions more likely
end of the warm-upWe now have every ingredient: information, gradients, vectors, images, sequences, and reward. Put them together and you get the architecture that changed everything.
introduces →reinforcement learningpolicyvalue functionexploration
← previousWords as Coordinatesnext →Attention Is All You Need