Training Deep Models: Optimization and Regularization
Making Training Work
Gradient Descent
Gradient descent updates parameters in the direction that reduces loss. With learning rate , the update is often written as .
Mini-Batch Training
Instead of computing gradients on the full dataset, mini-batch training uses small batches. This usually makes learning faster and more scalable.
Learning rate matters
If the learning rate is too large, training can diverge. If it is too small, training can become painfully slow.
What does gradient descent do?
Gradient descent uses the gradient to improve the model by reducing the loss.
Correct answer: Moves parameters toward lower loss
Why is mini-batch training useful?
Mini-batches are common because they are practical for large datasets and hardware accelerators.
Correct answer: It balances computational efficiency with noisy but useful gradient estimates.
Regularization
Regularization techniques reduce overfitting by limiting model complexity or adding constraints. Examples include penalties, dropout, early stopping, and data augmentation.
Which method is a form of regularization?
Early stopping halts training before the model overfits.
Correct answer: Early stopping
What is the purpose of regularization?
Regularization helps a model perform well on unseen data.
Correct answer: To improve generalization by discouraging overly complex models.