DA5453: Learning from Human Preferences
This course covers the fundamental concepts of learning human preferences from comparisons and choices, with modern applications, primarily LLM alignment. The course has three modules, listed below. Note that the course is mathematically heavy. Students taking this elective are expected to have a research mindset—to go deeper into the concepts covered in class, and to self-study topics they might be weak in.
Prerequisites: one basic ML course (DA5400 or equivalent); comfort with probability and optimisation (DA5000 or equivalent). Basic familiarity with LLMs and PyTorch would help. Multi-armed bandits, reinforcement learning, and LLM architecture details are not pre-requisites.
The course is appropriate for the following students:
- B.Tech: 3rd year and above (any department)
- IDDD Data Science: 4th year and above
- M.Tech: 2nd year and above
- Research scholars (MS/PhD)
Logistics
Slot J · Venue: NAC 2-633 · 9 Credits
| Slot | Day & time | Use |
|---|---|---|
| J1 | Wednesday 2:00–3:15 PM | Lecture |
| J2 | Thursday 3:30–4:45 PM | Lecture |
| J3 | Monday 5:00–5:50 PM | Tutorials, vivas |
First lecture: Wed 29 Jul 2026 · Last lecture: Thu 5 Nov 2026.
Modules
The course will consist of three modules, taught in the following order.
| Module | Topic | Lectures |
|---|---|---|
| M1 | Choice modeling foundations (Bradley-Terry and Multinomial Logit, estimation and guarantees, limitations) | 8 |
| M2 | LLM alignment–reinforcement learning from human feedback (RLHF) and Direct Preference Optimisation (DPO) | 6 |
| M3 | Online learning from comparisons (active learning, duelling bandits, contextual duelling bandits) | 6 |
Assessment
| Component | Marks | When |
|---|---|---|
| Quiz-1 (Module 1) | 20 | Thu 27 Aug, in class |
| Coding Assignment (Module 2) | 20 | Due Sun 4 Oct, 11:59 PM IST |
| Quiz-2 (Module 3) | 20 | Thu 15 Oct, in class |
| Project | 40 | Presentations: 21 Oct–5 Nov; report: 27 Nov |
Module 1 comes with an ungraded take-home assignment. Quizzes are closed-book and last 2 hours. Quiz-1 is based largely on the Module 1 take-home assignment.
The project is to be done in pairs: master a paper from a curated list, teach it to the class in ~15 minutes, then implement the paper’s main idea in code (or dive deeper into the theorems) and write a report. More instructions will be given in class.
No end-semester exam.
Materials
Module 1
Module 2
- Course notes: PDF
- Coding assignment: assignment page
- Submission form: submit your assignment — due Sunday, 4 October 2026, 11:59 PM IST
Module 3
- Course notes: PDF
Capstone paper list — suggested papers.
Suggested reading: Reinforcement Learning from Human Feedback, Nathan Lambert (2025); and Machine Learning from Human Preferences, Truong, Haupt and Koyejo (Stanford, 2025).
