Stanford Root

Schedule

Stanford Root

Schedule

CS 221M

Mechanistic Interpretability

UNITS:3
GRADING:Letter or Credit/No Credit
LEVEL:Graduate
GER:—

What is the internal structure of modern neural networks and how can we study it? This course provides a broad and deep introduction to interpretability, the subfield of machine learning concerned with understanding precisely how models process information and why they produce the outputs they do. We will cover topics such as probing, steering, causal abstraction, and sparse autoencoders, with a particular emphasis on causal methods and large language models. The course will include guest lectures from leading interpretability labs across academia and industry.

Syllabus for selected term:
View Spring 2027 Syllabus

Sections

1 Term
Lecture 1Open
ID: 6516
0 / 30 enrolled
DAYS:Monday, Wednesday
TIME:2:30 PM – 4:20 PM
LOCATION:TBD
INSTRUCTOR:
Icard, Thomas
3units

CS 221M: Mechanistic Interpretability

3 units · Letter or Credit/No Credit

What is the internal structure of modern neural networks and how can we study it? This course provides a broad and deep introduction to interpretability, the subfield of machine learning concerned with understanding precisely how models process information and why they produce the outputs they do. We will cover topics such as probing, steering, causal abstraction, and sparse autoencoders, with a particular emphasis on causal methods and large language models. The course will include guest lectures from leading interpretability labs across academia and industry.

Offered in Spring 2027 at Stanford University.

Spring 2027 sections

  • Lecture — Monday Wednesday 2:30 PM – 4:20 PM — Icard, Thomas (Graduate)

More CS courses

  • CS 212: Operating Systems and Systems Programming
  • CS 214: Selected Reading of Computer Science Research
  • CS 217: Hardware Accelerators for Machine Learning (EE 244)
  • CS 218: Information Integrity
  • CS 220: Researching, Presenting and Publishing Work in AI & Education (EDUC 481)
  • CS 221: Artificial Intelligence: Principles and Techniques
  • CS 223A: Introduction to Robotics (ME 320)
  • CS 224G: Apps With LLMs Inside
  • CS 224N: Natural Language Processing with Deep Learning (LINGUIST 284, SYMSYS 195N)
  • CS 224R: Deep Reinforcement Learning
  • CS 224U: Natural Language Understanding (LINGUIST 188, LINGUIST 288, SYMSYS 195U)
  • CS 224V: Agentic AI

All CS courses · All departments