Reinforcement Learning: In Theory and In Practice
Fall 2026. ESE 3990. Tue / Thu 10:15-11:45. DRLB 3C4

Announcements
Course Overview
This course provides a comprehensive treatment of reinforcement learning, bridging the gap from classical foundations in dynamic programming and optimal control to the cutting edge of modern AI. Students will explore how agents learn optimal behaviors in possibly uncertain and partially observable environments. We will start from simple MDP and Bandit games and gradually build towards high-dimensional state spaces using deep neural networks for function approximation. Towards the end of the course, we will cover case studies in autonomous robotics and the fine-tuning of large language models (RLHF). At the end of this course, students will be equipped to use RL solutions for advanced applications and critically evaluate their performance, advantages, and limitations against alternative optimization methodologies.
Prerequisites
This is a undergraduate-level course. Students are expected to have prior knowledge in linear algebra and geometry. Prior experience with deep learning will be beneficial but not necessary.
Schedule
Foundation
Foundations of Reinforcement Learning
- SEP 8
- Planning by Dynamic Programming (DP)
- Slides • Reading: S&B, Ch. 4.
- SEP 10
- DP (II) and Model-free Prediction
- Slides (DP), Slides (MfP) • Reading: S&B, Ch. 5-6.
- SEP 15
- Model-free Value Prediction (II)
- Slides (MfP) • Reading: S&B, Ch. 5-6.
- SEP 17
- Model-free Control
- Slides• Reading: S&B, Ch. 6
- SEP 22
- Second Quiz + Recitation
- SEP 24
- Off-Policy Learning
- Slides • Reading: S&B, Ch. 5-6
- SEP 29
- Value Function Approximation
- Slides• Reading: S&B, Ch. 9
- OCT 1
-
No Class Fall Break
- OCT 6
- Policy Gradients (I)
- Slides
- OCT 8
- Policy Gradients (II)
- OCT 13
- Model-based RL
- OCT 15
- Fundamentals Recitation (No Quiz)
- OCT 20
- Midterm
Algorithms. On-Policy, Off-Policy, and Model-Based.
- OCT 22
- Advanced Policy Gradients
- OCT 27
- Advanced Policy Gradients
- OCT 29
- Advanced Off-policy Learning
- Nov 3
- Advanced Off-policy Learning
- NOV 5
- Advanced Model-based RL
- NOV 10
- Advanced Model-based RL
- NOV 12
- Recitation and Quiz. Intermediate Project Presentation
Applications and Frontiers
- NOV 17
- Multi-Task and Meta RL
- NOV 19
- Inverse RL
- NOV 24
- Case Study: RL on LLMs
- NOV 26
-
No Class Thanksgiving
- DEC 1
- Case Study: RL in Robotics.
- DEC 3
- Final Project Presentation
Instructors

Teaching Assistants

Related Courses
Deep Reinforcement Learning: CMU version UC Berkeley version.
Reinforcement Learning , UCL. (The foundation part of our course is heavily based on this material).