Convolutional, Recurrent, and Sequence Models
Structured data modalities
Why architecture matters
Different data types have different symmetries and dependencies. Images exhibit locality and translation structure, sequences exhibit order, and language requires context-sensitive dependency modeling.
Major architectures
Convolutional neural networks
- Exploit local connectivity and weight sharing
- Strong for images and grids
Recurrent neural networks
- Model temporal recurrence
- Useful for sequences and time series
What is the key idea behind convolutional weight sharing?
Weight sharing enables translation equivariance and parameter efficiency.
Correct answer: The same filter is applied across locations
Why do transformers often outperform RNNs on long-range dependencies?
This shortens dependency paths and improves parallelism.
Correct answer: Self-attention allows direct interaction between distant tokens without step-by-step recurrence.
Attention mechanism
Attention computes weighted combinations of values using query-key similarity. In scaled dot-product attention, enabling content-based retrieval over context.
Sequence modeling insight
Architectures encode inductive biases: CNNs for locality, RNNs for causality, and attention for flexible context selection.
Which property is most associated with CNNs?
Convolutions detect patterns regardless of position, giving translation equivariance.
Correct answer: Translation equivariance
Training sequence models
-
1
Prepare tokenized or structured inputs.
-
2
Choose architecture-specific positional or temporal encoding.
-
3
Train with sequence-level or token-level objectives.
-
4
Evaluate on downstream tasks and long-context generalization.
What does self-attention compute between tokens?
Attention measures how much each token should attend to others in the sequence.
Correct answer: A context-dependent weighted interaction or relevance score.