Topics and Objectives
- Function approximation and basis expansions
- Regression, natural cubic, and smoothing splines
- Regularization, effective degrees of freedom, and GCV
- Additive models and tensor-product splines
- Positive-definite kernels and RKHS
- Kernel computation, kernel ridge regression, and kernel principal
component analysis
- Kernel mean embeddings and maximum mean discrepancy
- The representer theorem and support vector regression
Module Schedule
During this four-week module, we will cover the following potential
topics as time permits. You will complete one homework assignment, due
by Thursday of the third week. The fourth week will be a Present and
Challenge week.
Homework
- [Homework 1]
[Download R Markdown]
- Due September 10, 11:59 PM
CT
- Submit through Gradescope using
course ID
1370811 and entry code
XE3YJ6.
- You are only required to complete
five of the six questions. You may choose any five.
Presentation Session
General logistics and scoring are described on the Present and Challenge page.
Selected teams are required to present a modern paper from the list
below. Each team may use up to 12 minutes for its presentation and up to
10 minutes for Q&A. The specific time allocation may be shorter
depending on the number of teams presenting in the session.
The basic requirements are:
- Introduce the problem and the fundamental idea of the paper. Explain
why the method was innovative relative to the literature at the time and
what advantages it offers.
- Implement the method and compare it with appropriate baselines in a
simulation study to evaluate how well the method works.
- Discuss potential extensions or adaptations of the method for
today’s machine learning and data analysis problems.
Teams may also propose a paper of their own choice. The paper
should:
- Be modern, preferably published after 2000, and closely related to
RKHS, kernel methods, or their applications.
- Address a focused problem with a central idea that can be clearly
explained within the presentation time.
- Contain a method that can be reasonably implemented and compared
with appropriate baselines, with a workload comparable to a standard
homework assignment.
Here are some candidate papers:
- Williams, C. K. I., & Seeger, M. (2000). Using the Nyström
Method to Speed Up Kernel Machines. Advances in Neural Information
Processing Systems, 13, 682-688. [link]
- Lodhi, H., Saunders, C., Shawe-Taylor, J., Cristianini, N., &
Watkins, C. (2002). Text Classification using String Kernels. Journal of
Machine Learning Research, 2, 419-444. [link]
- Rahimi, A., & Recht, B. (2007). Random Features for Large-Scale
Kernel Machines. Advances in Neural Information Processing Systems, 20,
1177-1184. [link]
- Pan, S. J., Tsang, I. W., Kwok, J. T., & Yang, Q. (2011). Domain
Adaptation via Transfer Component Analysis. IEEE Transactions on Neural
Networks, 22(2), 199-210. [link]
- Shervashidze, N., Schweitzer, P., van Leeuwen, E. J., Mehlhorn, K.,
& Borgwardt, K. M. (2011). Weisfeiler-Lehman Graph Kernels. Journal
of Machine Learning Research, 12(77), 2539-2561. [link]
- Muandet, K., Fukumizu, K., Dinuzzo, F., & Schölkopf, B. (2012).
Learning from Distributions via Support Measure Machines. Advances in
Neural Information Processing Systems, 25, 10-18. [link]
- Yamada, M., Jitkrittum, W., Sigal, L., Xing, E. P., & Sugiyama,
M. (2014). High-Dimensional Feature Selection by Feature-Wise Kernelized
Lasso. Neural Computation, 26(1), 185-207. [link]
- Li, Y., Swersky, K., & Zemel, R. (2015). Generative Moment
Matching Networks. Proceedings of the 32nd International Conference on
Machine Learning, PMLR 37, 1718-1727. [link]
- Kandasamy, K., & Yu, Y. (2016). Additive Approximations in High
Dimensional Nonparametric Regression via the SALSA. Proceedings of the
33rd International Conference on Machine Learning, PMLR 48, 69-78. [link]
- Kim, B., Khanna, R., & Koyejo, O. O. (2016). Examples are not
enough, learn to criticize! Criticism for Interpretability. Advances in
Neural Information Processing Systems, 29, 2280-2288. [link]
- Yu, F. X., Suresh, A. T., Choromanski, K. M., Holtmann-Rice, D. N.,
& Kumar, S. (2016). Orthogonal Random Features. Advances in Neural
Information Processing Systems, 29, 1975-1983. [link]
- Liu, F., Xu, W., Lu, J., Zhang, G., Gretton, A., & Sutherland,
D. J. (2020). Learning Deep Kernels for Non-Parametric Two-Sample Tests.
Proceedings of the 37th International Conference on Machine Learning,
PMLR 119, 6316-6326. [link]