DA6002W: Online and Reinforcement Learning
May–June 2026.
This is an elective course for students in the Web-enabled M.Tech. in Data Science and AI, offered by IIT Madras. I taught the online learning half of the course for seven weeks. In this iteration, the course focused only on bandits. I tried to tailor the course to make it more suitable for working professionals.
Lecture slides
- Introduction to Online Learning
- Best-Arm Identification in Bandits
- Regret Minimisation in Bandits
- Upper Confidence Bound for Bandits
- UCB Recap and Thompson Sampling for Bandits
- Thompson Sampling for Bandits–continued
- Introduction to Contextual Bandits
- A Deep Dive Into Linear Bandits
- LinUCB and LinTS in Action
- Linear Bandits in Practice
- Production Realities in Bandits
- Off-Policy Evaluation in Contextual Bandits
- Beyond Single-Arm Actions
- Bandits Recap and Looking Ahead
