← Back to CoursesArtificial Intelligence: Advanced

Neuroanatomy Explorer

Drag to rotate · scroll to zoom · click regions to explore

View
Loading 3D model…

Click a region
to explore it

Memory Deck

Flip each card and rate whether you knew it. Your score is saved.

Term
Definition

Deck complete — score saved.

Match the Pairs

Match each term to its definition. Finish the board to earn your score.

All matched — score saved.

Concept Constellation

Every key idea in this course, mapped as an explorable 3D constellation. Drag to rotate, scroll to zoom, click a node.

Click a node to read its definition.

Reinforcement Learning and Sequential Decision-Making

Manual: General · Subject: Artificial Intelligence

Explore Markov decision processes, value functions, policy optimization, and temporal credit assignment.

Learning by interaction

Core formulation

Reinforcement learning studies agents that learn from rewards obtained by interacting with an environment. In an MDP, dynamics are defined by states, actions, transition probabilities, rewards, and discount factor γ\gamma.

MDP elements

State ss
Current situation description
Action aa
Choice made by the agent
Reward rr
Immediate scalar feedback
Policy π\pi
Mapping from states to actions or action distributions

What does the discount factor γ\gamma control?

What is the Bellman equation used for?

Key RL methods

Value-based methods

  • Learn state or action values
  • Example: Q-learning

Policy-based methods

  • Directly optimize policies
  • Useful for continuous actions

Exploration and exploitation

An RL agent must balance exploiting known high-reward actions with exploring uncertain alternatives. Efficient exploration remains a major research challenge, especially in sparse-reward or high-dimensional settings.

⚠️

Credit assignment

Learning which earlier actions caused later rewards is difficult because feedback may be delayed, noisy, and confounded by stochastic transitions.

Which method directly learns an action-value function?

Name one challenge unique to reinforcement learning.