Course content
- Foundations of Reinforcement Learning
- Markov Decision Processes and Value Functions
- Bellman Equations and Dynamic Programming
- Model-Free Learning
- Function Approximation
- Introduction to Deep RL and DQN
- Policy Gradients: REINFORCE and Actor-Critic
- Proximal Policy Optimization
- RLHF: Aligning Language Models with Human Feedback
- Verifiable Rewards and GRPO
- The Reward Signal Problem for Agents
- Agentic RL: Environments, Trajectories, and the Training
- How do AI teams use RL in production?
Published on Apr 25, 2026