The Reward Signal Problem for Agents
RL Part 11: From verifiable rewards to LLM-as-a-judge.
RL Part 11: From verifiable rewards to LLM-as-a-judge.
RL Part 10: Dropping two models from the four-model pipeline, and building rewards you can trust.
Used by top products, including Anthropic, Google, etc.
Part 9: From human preferences to a trained reward signal, and the four-model PPO pipeline.
RL Part 8: Trust regions, the clipped surrogate, and the workhorse of modern RL.
RL Part 7: Learning the policy directly, from REINFORCE to actor-critic.
RL Part 6: From linear features to neural networks, and the engineering choices that makes deep value-based RL possible.
RL Part 5: From tables to parameterized value functions.