
Charlie and the Reasoning Room
Book IV — Reinforcement learning, from bandits to reasoning models.
The fourth and last room. Reinforcement learning derived and coded from scratch: the agent-environment loop, dynamic programming, Monte Carlo and temporal difference, SARSA, Q-learning and DQN, policy gradients, TRPO and PPO, then RLHF, GRPO, DPO and agentic RL, and finally the production projects that put RL into tutors, world models, robots, cars and codebases.
Read on your Kindle
We'll send this whole book straight to your Kindle — it opens natively, so you can resize the text, read fully offline, and it remembers where you left off. Nothing to download or manage.
Sending to Kindle is for subscribers — subscribe to read the whole library on your Kindle.
00Into the Reasoning Room3 capsules
The only room in the factory where the machine is never shown a correct answer.
01The Reasoning RoomSTARTfree16 min02The agent-environment loopMDP🔒18 min03A short history of RLHIST🔒18 min01The Three Horsemen3 capsules
Three ways to estimate how good a situation is, distinguished by how much of the future you wait for.
04Dynamic programmingDP🔒18 min05Monte Carlo methodsMC🔒18 min06Temporal-difference learningTD🔒17 min02Value Machines3 capsules
Learn a number for every state-action pair, then act greedily with respect to it.
07SARSA and Q-learningQLEARN🔒17 min08Deep Q-networksDQN🔒17 min09Replay buffers and target networksSTAB🔒15 min03Policy Machines3 capsules
Skip the value function and change the behaviour directly.
10Policy gradients and REINFORCEPG🔒19 min11Baselines, advantage and actor-criticGAE🔒17 min12TRPO and PPOPPO🔒20 min04Aligning Language Models4 capsules
Point the same machinery at a language model and the reward becomes a human preference.
13RLHF from scratchRLHF🔒18 min14Reward models and their failure modesREWARD🔒17 min15GRPO and reasoning modelsGRPO🔒18 min16DPO and the direct methodsDPO🔒16 min05Production RL5 capsules
Six projects where the reward was real and the environment did not come from a benchmark.
17Agentic RLAGENT🔒20 min18The Socratic tutorTUTOR🔒16 min19World models and dreamingDREAM🔒18 min20RL for robotics and drivingROBOT🔒19 min21RL for software, and RL infrastructureINFRA🔒17 minRatings & reviews
No ratings yet. Yours would be the first.