Topics and Objectives
- Markov decision processes, the Markov property, and discounted
rewards
- Value functions, Q-functions, and Bellman equations
- Bellman operators, contraction mappings, and value iteration
- Policy evaluation, policy improvement, and model-based versus
model-free learning
- Fitted Q-iteration with linear, nonlinear, and forest-based function
approximation
- Policy gradients, occupancy measures, and on-policy versus
off-policy learning
- Maximum-entropy objectives, soft actor-critic, and behavior
regularization in offline reinforcement learning
Module Schedule
This module contains four instructional weeks, excluding Fall Break.
We may cover the following potential topics as time permits. The fourth
instructional week will be a Present and Challenge week. Detailed Fall
2026 lecture materials are TBA.
- Markov decision processes, value functions, and Bellman
equations
- Bellman operators and the contraction mapping theorem
- Value iteration, policy evaluation, and policy improvement
- Fitted Q-iteration and function approximation
- Policy gradient methods
- Soft actor-critic and maximum-entropy objectives
- Offline reinforcement learning and behavior regularization
Homework
Presentation Session
For this module’s presentation topics, you should select a paper on
reinforcement learning, sequential decision making, AI agents, or
another closely related topic.
Here are some candidate papers:
References
- Sutton, R. S., & Barto, A. G. (2018). Reinforcement Learning: An
Introduction. MIT Press.
- Reinforcement Learning: Theory and Algorithms by A.
Agarwal, K. Brantley, N. Jiang, S. M. Kakade, and W. Sun.
- Algorithms for Reinforcement Learning by C.
Szepesvári.
- Puterman, M. L. (1994). Markov Decision Processes: Discrete
Stochastic Dynamic Programming. Wiley.
- Ernst, D., Geurts, P., & Wehenkel, L. (2005). Tree-based batch
mode reinforcement learning. Journal of Machine Learning Research, 6,
503-556.
- Williams, R. J. (1992). Simple statistical gradient-following
algorithms for connectionist reinforcement learning. Machine Learning,
8, 229-256.
- Schulman, J., Levine, S., Abbeel, P., Jordan, M., & Moritz, P.
(2015). Trust region policy optimization. Proceedings of the 32nd
International Conference on Machine Learning, 1889-1897.
- Schulman, J., Wolski, F., Dhariwal, P., Radford, A., & Klimov,
O. (2017). Proximal policy optimization algorithms.
arXiv:1707.06347.
- Haarnoja, T., Zhou, A., Abbeel, P., & Levine, S. (2018). Soft
actor-critic: Off-policy maximum entropy deep reinforcement learning
with a stochastic actor. Proceedings of the 35th International
Conference on Machine Learning, 1861-1870.
- Kumar, A., Fu, J., Soh, M., Tucker, G., & Levine, S. (2019).
Stabilizing off-policy Q-learning via bootstrapping error reduction.
Advances in Neural Information Processing Systems, 32.
- Fujimoto, S., Meger, D., & Precup, D. (2019). Off-policy deep
reinforcement learning without exploration. Proceedings of the 36th
International Conference on Machine Learning, 2052-2062.
- Wu, Y., Tucker, G., & Nachum, O. (2019). Behavior regularized
offline reinforcement learning. arXiv:1911.11361.