← Back to CoursesArtificial Intelligence: Intermediate

Neuroanatomy Explorer

Drag to rotate · scroll to zoom · click regions to explore

View
Loading 3D model…

Click a region
to explore it

Memory Deck

Flip each card and rate whether you knew it. Your score is saved.

Term
Definition

Deck complete — score saved.

Match the Pairs

Match each term to its definition. Finish the board to earn your score.

All matched — score saved.

Concept Constellation

Every key idea in this course, mapped as an explorable 3D constellation. Drag to rotate, scroll to zoom, click a node.

Click a node to read its definition.

Reinforcement Learning and Decision Making

Manual: General · Subject: Artificial Intelligence

Understand how agents learn to act through reward, feedback, and exploration.

Learning by Acting

The RL Loop

In reinforcement learning, an agent observes a state ss, chooses an action aa, receives a reward rr, and transitions to a new state. The objective is to maximize long-term expected reward.

Exploration vs Exploitation

Agents must balance trying new actions to discover better strategies with choosing actions that already seem effective. This tradeoff is central to reinforcement learning.

RL Concepts

Policy
A mapping from states to actions.
Reward
Feedback signal used to evaluate actions.
Value function
Expected long-term return from a state or action.
Episode
A complete sequence of interaction steps.

What is the agent trying to maximize in reinforcement learning?

What is the exploration-exploitation tradeoff?

Supervised Learning vs Reinforcement Learning

Supervised Learning

  • Uses labeled examples
  • Learns a direct input-output mapping
  • Feedback is immediate and explicit

Reinforcement Learning

  • Uses reward signals
  • Learns through interaction
  • Feedback may be delayed

A policy in RL is:

Why can rewards be tricky to design?