← Back to CoursesAgentic AI: PhD Level

Neuroanatomy Explorer

Drag to rotate · scroll to zoom · click regions to explore

View
Loading 3D model…

Click a region
to explore it

Memory Deck

Flip each card and rate whether you knew it. Your score is saved.

Term
Definition

Deck complete — score saved.

Match the Pairs

Match each term to its definition. Finish the board to earn your score.

All matched — score saved.

The Agent Loop in 3D

Watch a thought travel through Perceive → Plan → Act → Observe. Drag to rotate, scroll to zoom, click a node.

Click a node to read its definition.

Foundations: Planning, Search, Reinforcement Learning, and Control

Manual: General · Subject: Agentic AI

Develop the algorithmic foundations of agentic behavior from classical planning to modern sequential decision-making.

Sequential decision-making

Classical view

Agentic AI inherits ideas from planning, search, and control. In the standard formalism, the agent chooses actions ata_t to maximize expected return E[∑t=0∞γtrt]\mathbb{E}[\sum_{t=0}^{\infty} \gamma^t r_t] under uncertainty.

Planning and control

Classical planners reason over symbolic states and transition models, whereas reinforcement learning learns policies from interaction. Modern agentic systems often hybridize both: explicit planning for structure, learning for flexibility.

Core concepts

State
A representation of the world relevant to action selection
Policy
A mapping from states or observations to actions
Value function
Expected long-term return from a state or action
Exploration
Behavior that gathers information for future benefit
Exploitation
Selecting actions that maximize current estimated utility

A planning loop

  1. 1

    Step 1: Formulate the goal and constraints.

  2. 2

    Step 2: Predict outcomes of candidate actions.

  3. 3

    Step 3: Search or optimize over action sequences.

  4. 4

    Step 4: Execute the best action.

  5. 5

    Step 5: Update beliefs from feedback and repeat.

Which term denotes the long-term desirability of a state under a policy?

Why are hybrid agent architectures common?

State the Bellman intuition in one phrase.

Name one limitation of pure planning in open-world agentic systems.