DA6002W: Online and Reinforcement Learning

May–June 2026.

This is an elective course for students in the Web-enabled M.Tech. in Data Science and AI, offered by IIT Madras. I taught the online learning half of the course for seven weeks. In this iteration, the course focused only on bandits. I tried to tailor the course to make it more suitable for working professionals.

Lecture slides

  1. Introduction to Online Learning
  2. Best-Arm Identification in Bandits
  3. Regret Minimisation in Bandits
  4. Upper Confidence Bound for Bandits
  5. UCB Recap and Thompson Sampling for Bandits
  6. Thompson Sampling for Bandits–continued
  7. Introduction to Contextual Bandits
  8. A Deep Dive Into Linear Bandits
  9. LinUCB and LinTS in Action
  10. Linear Bandits in Practice
  11. Production Realities in Bandits
  12. Off-Policy Evaluation in Contextual Bandits
  13. Beyond Single-Arm Actions
  14. Bandits Recap and Looking Ahead