Learning from Reward
Supervised learning needs labels. But how do you label the right move in a game you are still learning to play? Reinforcement learning (RL) works from reward alone, trial and error, over time. It is how agents learn to act.
The core objects
A policy is the agent’s rule for choosing actions — the thing being learned. A value function estimates how good a state is, or how good an action is in that state. The reward signal might be sparse: a single +1 at the end of a long game, with nothing to say which of the thousand moves was responsible. That is the credit-assignment problem again, now spread across time.
Exploration is the central tension. Take the action you think is best, and you never learn whether something better exists. Try new things, and you waste reward. Every RL algorithm is, in part, a policy for balancing the two.
The two great families
Q-learning learns the value of every action and picks the best. Policy gradient methods skip values and push probabilities directly toward actions that led to reward. Modern systems usually blend the two.