A graduate-level course focused on principles and techniques for designing and building large-scale disaggregated AI inference systems. The course explores architectural, algorithmic, system-level, and kernel implementation techniques for serving models with low latency, high throughput, and low cost on heterogeneous machines. Topics include model optimization techniques, distributed serving, AI hardware and performance modelling, kernel programming, and AI compilers.
3 units · Letter or Credit/No Credit
A graduate-level course focused on principles and techniques for designing and building large-scale disaggregated AI inference systems. The course explores architectural, algorithmic, system-level, and kernel implementation techniques for serving models with low latency, high throughput, and low cost on heterogeneous machines. Topics include model optimization techniques, distributed serving, AI hardware and performance modelling, kernel programming, and AI compilers.
Offered in Autumn 2026 at Stanford University.