← Back to CoursesAgentic AI: Advanced

Neuroanatomy Explorer

Drag to rotate · scroll to zoom · click regions to explore

View
Loading 3D model…

Click a region
to explore it

Memory Deck

Flip each card and rate whether you knew it. Your score is saved.

Term
Definition

Deck complete — score saved.

Match the Pairs

Match each term to its definition. Finish the board to earn your score.

All matched — score saved.

The Agent Loop in 3D

Watch a thought travel through Perceive → Plan → Act → Observe. Drag to rotate, scroll to zoom, click a node.

Click a node to read its definition.

Formal Models: Agents, Environments, and Objectives

Manual: General · Subject: Agentic AI

Develop the mathematical and computational foundations for modeling agentic systems.

Formalizing Agency

Agent-Environment Interaction

A canonical abstraction is the agent-environment loop. At time tt, the agent observes oto_t, updates internal state hth_t, chooses action ata_t, and receives reward or feedback rtr_t. The environment then transitions to a new state according to its dynamics.

Decision Objective

Many agentic systems optimize expected cumulative utility. In reinforcement learning, this is often written as J(π)=Eπ[∑t=0Tγtrt]J(\pi)=\mathbb{E}_{\pi}\left[\sum_{t=0}^{T} \gamma^t r_t\right] where π\pi is the policy and γ∈[0,1)\gamma \in [0,1) is a discount factor.

ℹ️

Partial Observability

Real-world agents rarely observe the full state of the world; they infer latent state from incomplete, noisy signals.

Model Components

State
A representation of the current situation.
Observation
What the agent directly perceives.
Action
A choice that affects the environment.
Reward
A scalar signal used for optimization.
Policy
A decision rule over states or observations.

In the standard RL objective, what does the discount factor γ\gamma control?

What is the main difference between a fully observed MDP and a POMDP?

From Policies to Plans

Agentic AI often blends reactive policies with deliberative planning. The system may use a policy π(at∣ht)\pi(a_t \mid h_t) for rapid action selection and a planner that searches over future trajectories to satisfy constraints or maximize long-term value.

Which formulation best captures an agent operating under partial observability?