AI Foundationspredict · compress · act

concepts → Field

Field

Reinforcement learning

timeless10 connections

learning what to do by trying, from delayed reward

Where it sits

The prism has six jobs across and eight layers down. Its primary cell isact × L6, with a noted second cell. Hatched cells cannot exist — a GPU does not learn, an institution does not infer.

act×L0 — its theory is a branch of foundations, not of any product

What must come first — and what it unlocks

Left to right is reading order, derived from the prerequisite_of edges. Nothing here is hand-ordered: the diagram is the graph.

RewardControl theoryMarkov decision processModel-based RLReinforcement learningReasoning modelRLHF

Before it: Reward · Control theory · Markov decision process · Model-based RL

Unlocks: Reasoning model · RLHF

What kind of thing it is — and what it is made of

It is made of: Policy, Value function.

How it goes wrong — and what answers that

known failure modes
Reward hackinggetting the reward without doing the thing — the optimiser finds the loophole

Where to read it

The chapter that introduces it, and any chapter that uses it again.

23Learning from Rewardact 5 · The Connectionist Turn

Where it comes from

Every connection

All 10 edges touching this node, grouped by relation family — the sections above are highlights from this list. Colours match the relation families inthe atlas.

Structure · 2
has as a partPolicyMechanism
has as a partValue functionEstimator
Order · 6
must come beforeReasoning modelArchitecture
must come beforeRLHFAlgorithm
requiresRewardObjective
requiresControl theoryField
requiresMarkov decision processMechanism
requiresModel-based RLMechanism
Flow · 1
is used byRLHFAlgorithm
Failure · 1
fails byReward hackingFailureMode

This page is a projection of one node in src/data/concepts.ts. It has no prose file of its own — 230 declared edges produce all 236 of these pages. Edit an edge and both endpoints change.