In this course, we discuss creative and impactful ways to reuse the same dataset multiple times, including techniques such as extracting random or systematic subsets, resampling, shuffling, and introducing noise to the original dataset. These can be used in many tasks, from calculating confidence intervals and p-values, to finding the influential data points within a dataset, to constructing improved models, and choosing parameters we don't know how to specify a priori. By the end of the course, the students will understand classical terms like bootstrap, cross validation, and permutation tests, as well as more recent terms like random forests, bagging, data models, and model collapse.
3 units · Letter or Credit/No Credit
In this course, we discuss creative and impactful ways to reuse the same dataset multiple times, including techniques such as extracting random or systematic subsets, resampling, shuffling, and introducing noise to the original dataset. These can be used in many tasks, from calculating confidence intervals and p-values, to finding the influential data points within a dataset, to constructing improved models, and choosing parameters we don't know how to specify a priori. By the end of the course, the students will understand classical terms like bootstrap, cross validation, and permutation tests, as well as more recent terms like random forests, bagging, data models, and model collapse.
Offered in Spring 2027 at Stanford University.