← Back to CoursesArtificial Intelligence: PhD Level

Neuroanatomy Explorer

Drag to rotate · scroll to zoom · click regions to explore

View
Loading 3D model…

Click a region
to explore it

Memory Deck

Flip each card and rate whether you knew it. Your score is saved.

Term
Definition

Deck complete — score saved.

Match the Pairs

Match each term to its definition. Finish the board to earn your score.

All matched — score saved.

Concept Constellation

Every key idea in this course, mapped as an explorable 3D constellation. Drag to rotate, scroll to zoom, click a node.

Click a node to read its definition.

Machine Learning Theory and Generalization

Manual: General · Subject: Artificial Intelligence

Studies supervised learning, statistical risk, bias-variance trade-offs, and theoretical foundations of generalization.

Learning from Data

Statistical learning view

Machine learning seeks functions that generalize from finite data to unseen inputs. The central objects are a hypothesis class, a loss function, and a data-generating distribution.

Risk minimization

The population risk is R(f)=E(x,y)∼D[ℓ(f(x),y)]R(f)=\mathbb{E}_{(x,y)\sim \mathcal{D}}[\ell(f(x),y)], while empirical risk uses a sample average. Learning theory asks when minimizing empirical risk also reduces population risk.

Foundational terms

Bias
Error from overly restrictive assumptions.
Variance
Sensitivity to data fluctuations.
Overfitting
Excellent training performance but poor generalization.
Regularization
Penalization or constraint to improve generalization.

Major learning paradigms

ParadigmObjectiveTypical challenge
SupervisedPredict labels or targetsLabel scarcity and noise
UnsupervisedDiscover structureEvaluation ambiguity
Self-supervisedLearn from surrogate tasksObjective design
Semi-supervisedUse few labels with many unlabeled examplesConfirmation bias

What does regularization primarily help control?

Why can a model with lower training loss still have worse test performance?

Bias-variance perspectives

High-bias model

  • Simpler decision boundary
  • May underfit
  • Often stable

High-variance model

  • More flexible boundary
  • May overfit
  • Sensitive to data
🔑

Open problem

Explaining why overparameterized models generalize so well remains a major theoretical challenge in modern AI.