← Back to CoursesArtificial Intelligence: PhD Level

Neuroanatomy Explorer

Drag to rotate · scroll to zoom · click regions to explore

View
Loading 3D model…

Click a region
to explore it

Memory Deck

Flip each card and rate whether you knew it. Your score is saved.

Term
Definition

Deck complete — score saved.

Match the Pairs

Match each term to its definition. Finish the board to earn your score.

All matched — score saved.

Concept Constellation

Every key idea in this course, mapped as an explorable 3D constellation. Drag to rotate, scroll to zoom, click a node.

Click a node to read its definition.

Reinforcement Learning and Sequential Decision-Making

Manual: General · Subject: Artificial Intelligence

Introduces Markov decision processes, dynamic programming, policy optimization, and modern RL algorithms.

Learning by Interaction

Sequential decision-making

Reinforcement learning studies agents that learn through interaction, balancing immediate rewards and long-term return. The canonical formalism is the Markov decision process, defined by states, actions, transitions, rewards, and discounting.

Value functions

The state-value function is Vπ(s)=Eπ[∑t=0∞γtrt∣s0=s]V^\pi(s)=\mathbb{E}_\pi\left[\sum_{t=0}^\infty \gamma^t r_t \mid s_0=s\right], and the action-value function is Qπ(s,a)Q^\pi(s,a). These quantities support policy evaluation and control.

RL algorithm families

FamilyExampleStrength
Dynamic programmingValue iterationExact if model is known
Monte Carlo methodsPolicy evaluation from episodesSimple and model-free
Temporal-difference learningQ-learningBootstraps from partial returns
Policy gradientsREINFORCE, actor-criticDirectly optimizes stochastic policies
⚠️

RL is data hungry

Many RL algorithms are sample-inefficient, sensitive to reward design, and unstable under function approximation; improving these properties is a central research challenge.

Which statement best describes the exploration-exploitation trade-off?

Why is off-policy learning useful?

Policy optimization styles

Value-based

  • Learn action values
  • Derive policy from values
  • Can struggle with continuous actions

Policy-based

  • Directly parameterize the policy
  • Works naturally in continuous spaces
  • Often higher gradient variance

Designing an RL experiment

  1. 1

    Step 1: Define reward, termination, and discounting.

  2. 2

    Step 2: Choose whether a model is available.

  3. 3

    Step 3: Select an algorithm matched to action-space structure.

  4. 4

    Step 4: Measure sample efficiency and stability.

  5. 5

    Step 5: Evaluate robustness across seeds and environment variations.