← Back to CoursesArtificial Intelligence: Advanced

Neuroanatomy Explorer

Drag to rotate · scroll to zoom · click regions to explore

View
Loading 3D model…

Click a region
to explore it

Memory Deck

Flip each card and rate whether you knew it. Your score is saved.

Term
Definition

Deck complete — score saved.

Match the Pairs

Match each term to its definition. Finish the board to earn your score.

All matched — score saved.

Concept Constellation

Every key idea in this course, mapped as an explorable 3D constellation. Drag to rotate, scroll to zoom, click a node.

Click a node to read its definition.

Convolutional, Recurrent, and Sequence Models

Manual: General · Subject: Artificial Intelligence

Study specialized architectures for spatial and sequential data, including CNNs, RNNs, and attention-based alternatives.

Structured data modalities

Why architecture matters

Different data types have different symmetries and dependencies. Images exhibit locality and translation structure, sequences exhibit order, and language requires context-sensitive dependency modeling.

Major architectures

Convolutional neural networks

  • Exploit local connectivity and weight sharing
  • Strong for images and grids

Recurrent neural networks

  • Model temporal recurrence
  • Useful for sequences and time series

What is the key idea behind convolutional weight sharing?

Why do transformers often outperform RNNs on long-range dependencies?

Attention mechanism

Attention computes weighted combinations of values using query-key similarity. In scaled dot-product attention, Attention(Q,K,V)=softmax(QKTdk)V\text{Attention}(Q,K,V)=\text{softmax}\left(\frac{QK^T}{\sqrt{d_k}}\right)V enabling content-based retrieval over context.

ℹ️

Sequence modeling insight

Architectures encode inductive biases: CNNs for locality, RNNs for causality, and attention for flexible context selection.

Which property is most associated with CNNs?

Training sequence models

  1. 1

    Prepare tokenized or structured inputs.

  2. 2

    Choose architecture-specific positional or temporal encoding.

  3. 3

    Train with sequence-level or token-level objectives.

  4. 4

    Evaluate on downstream tasks and long-context generalization.

What does self-attention compute between tokens?