AI Foundationspredict · compress · act

concepts → Algorithm

Algorithm

RLHF

also called reinforcement learning from human feedback

decade7 connections

tilt a pretrained model toward human preference — how raw prediction becomes behaviour

the formal statement
max_π E[r] − β·KL(π ‖ π_ref)

Where it sits

The prism has six jobs across and eight layers down. Its primary cell islearn × L4, with a noted second cell. Hatched cells cannot exist — a GPU does not learn, an institution does not infer.

Overlays: Safety & alignment. act×L4 — it is a policy-optimisation method, not just supervised fitting

What must come first — and what it unlocks

Left to right is reading order, derived from the prerequisite_of edges. Nothing here is hand-ordered: the diagram is the graph.

Fine-tuningReinforcement learningRLHFGRPO

Before it: Fine-tuning · Reinforcement learning

Unlocks: GRPO

What kind of thing it is — and what it is made of

It is made of: Reward model.

Where to read it

The chapter that introduces it, and any chapter that uses it again.

28Teaching Tasteact 6 · The Attention Revolution

Where it comes from

Every connection

All 7 edges touching this node, grouped by relation family — the sections above are highlights from this list. Colours match the relation families inthe atlas.

Structure · 1
has as a partReward modelModel
Order · 3
requiresFine-tuningAlgorithm
must come beforeGRPOAlgorithm
Flow · 1
Lineage · 2
is improved byGRPOAlgorithm
is improved byDirect preference optimisationAlgorithm

This page is a projection of one node in src/data/concepts.ts. It has no prose file of its own — 230 declared edges produce all 236 of these pages. Edit an edge and both endpoints change.