DA5453: Learning from Human Preferences

This course covers the fundamental concepts of learning human preferences from comparisons and choices, with modern applications, primarily LLM alignment. The course has three modules, listed below. Note that the course is mathematically heavy. Students taking this elective are expected to have a research mindset—to go deeper into the concepts covered in class, and to self-study topics they might be weak in.

Prerequisites: one basic ML course (DA5400 or equivalent); comfort with probability and optimisation (DA5000 or equivalent). Basic familiarity with LLMs and PyTorch would help. Multi-armed bandits, reinforcement learning, and LLM architecture details are not pre-requisites.

The course is appropriate for the following students:

  • B.Tech: 3rd year and above (any department)
  • IDDD Data Science: 4th year and above
  • M.Tech: 2nd year and above
  • Research scholars (MS/PhD)

Logistics

Slot J · Venue: NAC 2-633 · 9 Credits

SlotDay & timeUse
J1Wednesday 2:00–3:15 PMLecture
J2Thursday 3:30–4:45 PMLecture
J3Monday 5:00–5:50 PMTutorials, vivas

First lecture: Wed 29 Jul 2026 · Last lecture: Thu 5 Nov 2026.

Modules

The course will consist of three modules, taught in the following order.

ModuleTopicLectures
M1Choice modeling foundations (Bradley-Terry and Multinomial Logit, estimation and guarantees, limitations)8
M2LLM alignment–reinforcement learning from human feedback (RLHF) and Direct Preference Optimisation (DPO)6
M3Online learning from comparisons (active learning, duelling bandits, contextual duelling bandits)6

Assessment

ComponentMarksWhen
Quiz-1 (Module 1)20Thu 27 Aug, in class
Coding Assignment (Module 2)20Due Sun 4 Oct, 11:59 PM IST
Quiz-2 (Module 3)20Thu 15 Oct, in class
Project40Presentations: 21 Oct–5 Nov; report: 27 Nov

Module 1 comes with an ungraded take-home assignment. Quizzes are closed-book and last 2 hours. Quiz-1 is based largely on the Module 1 take-home assignment.

The project is to be done in pairs: master a paper from a curated list, teach it to the class in ~15 minutes, then implement the paper’s main idea in code (or dive deeper into the theorems) and write a report. More instructions will be given in class.

No end-semester exam.

Materials

Module 1

  • Course notes: PDF
  • Take-home assignment: PDF (ungraded)

Module 2

Module 3

  • Course notes: PDF

Capstone paper list — suggested papers.

Suggested reading: Reinforcement Learning from Human Feedback, Nathan Lambert (2025); and Machine Learning from Human Preferences, Truong, Haupt and Koyejo (Stanford, 2025).