Skip to main content

Reinforcement Learning Course

A series of technical deep dives on Reinforcement Learning that covers fundamentals and background, the classical techniques, MDPs, Bellman equations, deep RL methods, how RL is used to train modern language models, agentic RL, and much more.

Start Course

Course content

  1. Foundations of Reinforcement Learning
  2. Markov Decision Processes and Value Functions
  3. Bellman Equations and Dynamic Programming
  4. Model-Free Learning
  5. Function Approximation
  6. Introduction to Deep RL and DQN
  7. Policy Gradients: REINFORCE and Actor-Critic
  8. Proximal Policy Optimization
  9. RLHF: Aligning Language Models with Human Feedback
  10. Verifiable Rewards and GRPO
  11. The Reward Signal Problem for Agents
  12. Agentic RL: Environments, Trajectories, and the Training
  13. How do AI teams use RL in production?
Published on Apr 25, 2026