
Reinforcement Learning
Build every RL algorithm by hand.
From the agent-environment loop to DQN, policy gradients, RLHF and GRPO, every reinforcement learning algorithm is derived and coded from scratch. Watch agents learn from experience — landing a lunar lander, playing Atari, aligning a language model, and reasoning like DeepSeek-R1.
Read on your Kindle
We'll send this whole book straight to your Kindle — it opens natively, so you can resize the text, read fully offline, and it remembers where you left off. Nothing to download or manage.
Sending to Kindle is for subscribers — subscribe to read the whole library on your Kindle.
00Foundations of Reinforcement Learning5 capsules
Foundations of Reinforcement Learning — 5 chapters.
01What Is Reinforcement Learning?conceptfree12 min02The Agent-Environment Interfaceconcept🔒12 min03The Markov Property and MDPsmath🔒13 min04Rewards, Returns and Discountingmath🔒14 min05OpenAI Gymnasium: Your First Environmentcode🔒14 min01The Three Pillars of Classical RL6 capsules
The Three Pillars of Classical RL — 6 chapters.
06Value Functions and the Bellman Equationsmath🔒13 min07Dynamic Programming: Policy and Value Iterationconcept🔒13 min08Monte Carlo Methods: Learning from Episodesconcept🔒12 min09Temporal Difference Learningconcept🔒12 min10Q-Learning and SARSAmath🔒14 min11Project: Landing a Lunar Landerproject🔒13 min02Deep Q-Networks5 capsules
Deep Q-Networks — 5 chapters.
12The Birth of Deep RL: From Pong to Atariconcept🔒12 min13From Q-Tables to Q-Networksconcept🔒13 min14Experience Replay and Target Networksdeep-dive🔒14 min15The DQN Training Loopmath🔒14 min16Project: Build a DQN from Scratchproject🔒13 min03Policy Gradient Methods from Scratch6 capsules
Policy Gradient Methods from Scratch — 6 chapters.
17Why Optimize the Policy Directly?concept🔒13 min18The Policy Gradient Theoremmath🔒12 min19The REINFORCE Algorithmcode🔒12 min20REINFORCE with a Baselinemath🔒13 min21Advantage Functions and Actor-Criticconcept🔒13 min22Generalized Advantage Estimationdeep-dive🔒14 min04RLHF from Scratch7 capsules
RLHF from Scratch — 7 chapters.
23Why Language Models Need Alignmentconcept🔒12 min24The Three-Stage RLHF Pipelineconcept🔒14 min25Training a Reward Model from Preferencesmath🔒14 min26Project: An SLM That Writes Positive Storiesproject🔒13 min27The PPO Algorithm Explainedmath🔒12 min28The PPO Training Loop for Language Modelsdeep-dive🔒13 min29Project: A Reddit Post Summarizerproject🔒14 min05Build a Reasoning Model with GRPO4 capsules
Build a Reasoning Model with GRPO — 4 chapters.
30From PPO to GRPOconcept🔒12 min31How GRPO Worksmath🔒13 min32GRPO and the DeepSeek-R1 Revolutiondeep-dive🔒14 min33Project: Turn an SLM into a Reasoning Modelproject🔒14 min06Agentic Reinforcement Learning3 capsules
Agentic Reinforcement Learning — 3 chapters.
34What Is Agentic RL?concept🔒13 min35The Agentic RL Landscapedeep-dive🔒13 min36Project: An Agentic RAG App Trained with RLproject🔒13 min07From Learner to Researcher2 capsules
From Learner to Researcher — 2 chapters.
37Putting It All Togetherconcept🔒14 min38Doing Impactful RL Researchdeep-dive🔒13 minRatings & reviews
No ratings yet. Yours would be the first.