Vizuara Books
Charlie and the Reasoning Room
Free preview available. Sign in and subscribe to unlock the full book.
Vizuara AI Labs · advanced · Charlie and the Intelligence Factory

Charlie and the Reasoning Room

Book IV — Reinforcement learning, from bandits to reasoning models.

The fourth and last room. Reinforcement learning derived and coded from scratch: the agent-environment loop, dynamic programming, Monte Carlo and temporal difference, SARSA, Q-learning and DQN, policy gradients, TRPO and PPO, then RLHF, GRPO, DPO and agentic RL, and finally the production projects that put RL into tutors, world models, robots, cars and codebases.

advancedrlppogrporlhfreasoningagents
21 capsules123 figures~6 hoursby Dr. Raj Dandekar, Dr. Sreedath Panat

Read on your Kindle

We'll send this whole book straight to your Kindle — it opens natively, so you can resize the text, read fully offline, and it remembers where you left off. Nothing to download or manage.

Sending to Kindle is for subscribers — subscribe to read the whole library on your Kindle.

00Into the Reasoning Room3 capsules

The only room in the factory where the machine is never shown a correct answer.

01The Reasoning RoomSTARTfree16 min02The agent-environment loopMDP🔒18 min03A short history of RLHIST🔒18 min
01The Three Horsemen3 capsules

Three ways to estimate how good a situation is, distinguished by how much of the future you wait for.

04Dynamic programmingDP🔒18 min05Monte Carlo methodsMC🔒18 min06Temporal-difference learningTD🔒17 min
02Value Machines3 capsules

Learn a number for every state-action pair, then act greedily with respect to it.

07SARSA and Q-learningQLEARN🔒17 min08Deep Q-networksDQN🔒17 min09Replay buffers and target networksSTAB🔒15 min
03Policy Machines3 capsules

Skip the value function and change the behaviour directly.

10Policy gradients and REINFORCEPG🔒19 min11Baselines, advantage and actor-criticGAE🔒17 min12TRPO and PPOPPO🔒20 min
04Aligning Language Models4 capsules

Point the same machinery at a language model and the reward becomes a human preference.

13RLHF from scratchRLHF🔒18 min14Reward models and their failure modesREWARD🔒17 min15GRPO and reasoning modelsGRPO🔒18 min16DPO and the direct methodsDPO🔒16 min
05Production RL5 capsules

Six projects where the reward was real and the environment did not come from a benchmark.

17Agentic RLAGENT🔒20 min18The Socratic tutorTUTOR🔒16 min19World models and dreamingDREAM🔒18 min20RL for robotics and drivingROBOT🔒19 min21RL for software, and RL infrastructureINFRA🔒17 min

Ratings & reviews

No ratings yet. Yours would be the first.

Sign in to rate this book