Reinforcement Learning: In Theory and In Practice

Fall 2026. ESE 3990. Tue / Thu 10:15-11:45. DRLB 3C4

Image

Announcements

Course Overview

This course provides a comprehensive treatment of reinforcement learning, bridging the gap from classical foundations in dynamic programming and optimal control to the cutting edge of modern AI. Students will explore how agents learn optimal behaviors in possibly uncertain and partially observable environments. We will start from simple MDP and Bandit games and gradually build towards high-dimensional state spaces using deep neural networks for function approximation. Towards the end of the course, we will cover case studies in autonomous robotics and the fine-tuning of large language models (RLHF). At the end of this course, students will be equipped to use RL solutions for advanced applications and critically evaluate their performance, advantages, and limitations against alternative optimization methodologies.

Prerequisites

This is a undergraduate-level course. Students are expected to have prior knowledge in linear algebra and geometry. Prior experience with deep learning will be beneficial but not necessary.

Schedule

Foundation

AUG 25
Logistics and probability review
Slides • Notes
AUG 27
RL basic definitions
Slides • Reading: S&B, Ch. 1–2.
SEP 1
Markov Decision Processes
Slides • Reading: S&B, Ch. 3.
SEP 3
Recitation and Quiz.
Slides

Foundations of Reinforcement Learning

SEP 8
Planning by Dynamic Programming (DP)
Slides • Reading: S&B, Ch. 4.
SEP 10
DP (II) and Model-free Prediction
Slides (DP), Slides (MfP) • Reading: S&B, Ch. 5-6.
SEP 15
Model-free Value Prediction (II)
Slides (MfP) • Reading: S&B, Ch. 5-6.
SEP 17
Model-free Control
Slides• Reading: S&B, Ch. 6
SEP 22
Second Quiz + Recitation
SEP 24
Off-Policy Learning
Slides • Reading: S&B, Ch. 5-6
SEP 29
Value Function Approximation
Slides• Reading: S&B, Ch. 9
OCT 1
No Class Fall Break :fallen_leaf:
OCT 6
Policy Gradients (I)
Slides
OCT 8
Policy Gradients (II)
OCT 13
Model-based RL
OCT 15
Fundamentals Recitation (No Quiz)
OCT 20
Midterm

Algorithms. On-Policy, Off-Policy, and Model-Based.

OCT 22
Advanced Policy Gradients
OCT 27
Advanced Policy Gradients
OCT 29
Advanced Off-policy Learning
Nov 3
Advanced Off-policy Learning
NOV 5
Advanced Model-based RL
NOV 10
Advanced Model-based RL
NOV 12
Recitation and Quiz. Intermediate Project Presentation

Applications and Frontiers

NOV 17
Multi-Task and Meta RL
NOV 19
Inverse RL
NOV 24
Case Study: RL on LLMs
NOV 26
No Class Thanksgiving :turkey:
DEC 1
Case Study: RL in Robotics.
DEC 3
Final Project Presentation

Instructors

Avatar

Teaching Assistants

Avatar

Deep Reinforcement Learning: CMU version UC Berkeley version.

Reinforcement Learning , UCL. (The foundation part of our course is heavily based on this material).