Neural Networks and Deep Representation Learning
From features to representations
Neural network basics
A neural network composes linear transformations and nonlinear activations to learn hierarchical representations. Given layers , the model can represent highly complex functions.
Why are nonlinear activations necessary in deep networks?
Without nonlinearities, a stack of linear layers collapses to a single linear transformation.
Correct answer: To allow composition of linear models into a more expressive function class
Backpropagation in brief
-
1
Compute the forward pass and loss.
-
2
Apply the chain rule from output to input.
-
3
Accumulate gradients layer by layer.
-
4
Update parameters using an optimizer such as SGD or Adam.
What does backpropagation compute?
These gradients guide optimization by indicating how to change weights to reduce loss.
Correct answer: Gradients of the loss with respect to model parameters.
Optimization pathologies
Deep learning training can suffer from vanishing and exploding gradients, saddle points, poor conditioning, and sharp minima. Architectural choices, normalization, initialization, and optimization algorithms mitigate these issues.
Architectural motifs
MLP
- General-purpose dense layers
- Useful for tabular and latent features
Residual network
- Skip connections improve gradient flow
- Enables deeper architectures
Representation learning
Deep networks are powerful because they learn internal features adapted to the task, often outperforming handcrafted pipelines when data and compute are sufficient.
What is a common benefit of residual connections?
Skip connections help gradients propagate and make very deep models easier to train.
Correct answer: They help optimization in very deep networks
State one reason deep models can outperform shallow models.
Depth enables the reuse and composition of intermediate features.
Correct answer: They can learn hierarchical representations with greater compositional expressivity.