
RLHF from Scratch
Align a language model with human preferences, by hand.
Build the full RLHF stack from first principles: reward models from human preference pairs, policy gradients and the advantage, and PPO with importance sampling and clipping. Put it to work on two projects — a Positive TinyStories SLM and a Reddit post summarizer — then see how DPO simplifies the whole pipeline.
Read on your Kindle
We'll send this whole book straight to your Kindle — it opens natively, so you can resize the text, read fully offline, and it remembers where you left off. Nothing to download or manage.
Sending to Kindle is for subscribers — subscribe to read the whole library on your Kindle.
00Why Alignment4 capsules
Why Alignment — 4 chapters.
01What is RLHF and Why It Mattersconceptfree13 min02The Three-Stage Alignment Pipelineconcept🔒14 min03What You Will Buildconcept🔒13 min04Prerequisites and Notationmath🔒12 min01The LLM as an RL Agent5 capsules
The LLM as an RL Agent — 5 chapters.
05The Journey of a Token Through an LLMintuition🔒13 min06Tokenization and Embeddings Refresherconcept🔒14 min07The Agent-Environment Interfaceconcept🔒12 min08States, Actions, and the Policymath🔒15 min09Generation as a Trajectoryintuition🔒13 min02Reward Models6 capsules
Reward Models — 6 chapters.
10The Problem of Subjective Rewardsconcept🔒13 min11Preferences Instead of Scoresintuition🔒12 min12The Bradley-Terry Preference Modelmath🔒13 min13Building a Reward Head on an LLMcode🔒15 min14Training the Reward Modelcode🔒14 min15Evaluating and Trusting the Reward Modeldeep-dive🔒14 min03Policy Gradients and Advantage6 capsules
Policy Gradients and Advantage — 6 chapters.
16The Policy Gradient Objectivemath🔒13 min17REINFORCE and the Log-Derivative Trickmath🔒14 min18Baselines and Variance Reductionintuition🔒14 min19Value Functions and the Advantagemath🔒14 min20Generalized Advantage Estimation (GAE)math🔒13 min21Computing Advantages Over a Sequencecode🔒13 min04Proximal Policy Optimization6 capsules
Proximal Policy Optimization — 6 chapters.
22Why Vanilla Policy Gradients Are Unstableconcept🔒13 min23Importance Sampling for Off-Policy Updatesmath🔒13 min24The Clipped Surrogate Objectivemath🔒14 min25The KL Penalty and the Reference Modelconcept🔒17 min26The Actor-Critic Value Modelcode🔒15 min27The Full PPO Training Loopcode🔒15 min05Building It From Scratch6 capsules
Building It From Scratch — 6 chapters.
28Project A: A Positive TinyStories SLMproject🔒13 min29Training a Tiny Supervised Language Modelcode🔒14 min30Aligning the SLM with a Reward Signalproject🔒16 min31Project B: A Reddit Post Summarizerproject🔒13 min32Training the Summarizer's Reward Modelcode🔒16 min33PPO Fine-Tuning with OpenRLHFcode🔒13 min06Beyond PPO4 capsules
Beyond PPO — 4 chapters.
34Does RLHF Really Need a Reward Model?deep-dive🔒14 min35Direct Preference Optimization (DPO)math🔒15 min36PPO vs DPO: Tradeoffsdeep-dive🔒15 min37Putting It All Togetherconcept🔒12 minRatings & reviews
No ratings yet. Yours would be the first.