← Back to CoursesArtificial Intelligence: Intermediate

Neuroanatomy Explorer

Drag to rotate · scroll to zoom · click regions to explore

View
Loading 3D model…

Click a region
to explore it

Memory Deck

Flip each card and rate whether you knew it. Your score is saved.

Term
Definition

Deck complete — score saved.

Match the Pairs

Match each term to its definition. Finish the board to earn your score.

All matched — score saved.

Concept Constellation

Every key idea in this course, mapped as an explorable 3D constellation. Drag to rotate, scroll to zoom, click a node.

Click a node to read its definition.

Training Deep Models: Optimization and Regularization

Manual: General · Subject: Artificial Intelligence

Learn how deep learning models are trained stably and prevented from overfitting.

Making Training Work

Gradient Descent

Gradient descent updates parameters in the direction that reduces loss. With learning rate η\eta, the update is often written as θ←θ−η∇θL\theta \leftarrow \theta - \eta \nabla_\theta L.

Mini-Batch Training

Instead of computing gradients on the full dataset, mini-batch training uses small batches. This usually makes learning faster and more scalable.

💡

Learning rate matters

If the learning rate is too large, training can diverge. If it is too small, training can become painfully slow.

What does gradient descent do?

Why is mini-batch training useful?

Regularization

Regularization techniques reduce overfitting by limiting model complexity or adding constraints. Examples include L2L_2 penalties, dropout, early stopping, and data augmentation.

Which method is a form of regularization?

What is the purpose of regularization?