Stanford Root

Schedule

Stanford Root

Schedule

CS 349M

Machine Learning for Software Engineering

UNITS:3-4
GRADING:Letter or Credit/No Credit
LEVEL:Graduate
GER:—

In recent years, tools based on machine learning have become increasingly prevalent in the software engineering field. The ubiquity of machine learning is an important factor, but just as important is the availability of software engineering data: there are billions of lines of code available in public repositories (e.g. on GitHub), there is the change history of that code, there are discussion fora (e.g. Stack Overflow) that contain a wealth of information for developers, companies have access to telemetry on their apps from millions of users, and so on. The scale of software engineering data has permitted machine learning and statistical approaches to imagine tools that are beyond the capabilities of traditional, semantics-based approaches. In this graduate seminar, students will learn the various ways in which code and related artifacts can be treated as data, and how various developer tools can be built by applying machine learning over this data. The course will consist of discussion of a selection of research papers, as well as a hands-on project that can be done in small groups. Prerequisites: Familiarity with basic machine learning, and either CS 143 or CS 295.

Syllabus for selected term:
View Spring 2027 Syllabus

Sections

1 Term
Lecture 1Open
ID: 26171
0 / 999 enrolled
DAYS:TBD
TIME:TBD
LOCATION:TBD
units

CS 349M: Machine Learning for Software Engineering

3-4 units · Letter or Credit/No Credit

In recent years, tools based on machine learning have become increasingly prevalent in the software engineering field. The ubiquity of machine learning is an important factor, but just as important is the availability of software engineering data: there are billions of lines of code available in public repositories (e.g. on GitHub), there is the change history of that code, there are discussion fora (e.g. Stack Overflow) that contain a wealth of information for developers, companies have access to telemetry on their apps from millions of users, and so on. The scale of software engineering data has permitted machine learning and statistical approaches to imagine tools that are beyond the capabilities of traditional, semantics-based approaches. In this graduate seminar, students will learn the various ways in which code and related artifacts can be treated as data, and how various developer tools can be built by applying machine learning over this data. The course will consist of discussion of a selection of research papers, as well as a hands-on project that can be done in small groups. Prerequisites: Familiarity with basic machine learning, and either CS143 or CS295.

Offered in Spring 2027 at Stanford University.

Spring 2027 sections

  • Lecture — TBA TBA (Graduate)

More CS courses

  • CS 348K: Visual Computing Systems
  • CS 348N: Neural Models for 3D Geometry
  • CS 349D: AI Inference Infrastructure
  • CS 349E: Efficient ML Infrastructure at Scale
  • CS 349F: Fabric Architectures For AI Systems
  • CS 349H: Software Techniques for Emerging Hardware Platforms (EE 349)
  • CS 350S: Privacy-Preserving Systems
  • CS 353: Seminar on Logic & Formal Philosophy (PHIL 391)
  • CS 354: Topics in Intractability: Unfulfilled Algorithmic Fantasies
  • CS 355: Advanced Topics in Cryptography
  • CS 356: Topics in Computer and Network Security
  • CS 357S: Formal Methods for Computer Systems

All CS courses · All departments