Machine Learning Theory and Generalization
Learning from Data
Statistical learning view
Machine learning seeks functions that generalize from finite data to unseen inputs. The central objects are a hypothesis class, a loss function, and a data-generating distribution.
Risk minimization
The population risk is , while empirical risk uses a sample average. Learning theory asks when minimizing empirical risk also reduces population risk.
Foundational terms
Major learning paradigms
| Paradigm | Objective | Typical challenge |
|---|---|---|
| Supervised | Predict labels or targets | Label scarcity and noise |
| Unsupervised | Discover structure | Evaluation ambiguity |
| Self-supervised | Learn from surrogate tasks | Objective design |
| Semi-supervised | Use few labels with many unlabeled examples | Confirmation bias |
What does regularization primarily help control?
Regularization constrains or penalizes overly flexible models to improve generalization.
Correct answer: Model complexity and overfitting
Why can a model with lower training loss still have worse test performance?
Training loss measures fit to the sample, not the underlying distribution.
Correct answer: Because it may overfit the training data and learn spurious patterns that do not generalize.
Bias-variance perspectives
High-bias model
- Simpler decision boundary
- May underfit
- Often stable
High-variance model
- More flexible boundary
- May overfit
- Sensitive to data
Open problem
Explaining why overparameterized models generalize so well remains a major theoretical challenge in modern AI.