Formal Models: Agents, Environments, and Objectives
Formalizing Agency
Agent-Environment Interaction
A canonical abstraction is the agent-environment loop. At time , the agent observes , updates internal state , chooses action , and receives reward or feedback . The environment then transitions to a new state according to its dynamics.
Decision Objective
Many agentic systems optimize expected cumulative utility. In reinforcement learning, this is often written as where is the policy and is a discount factor.
Partial Observability
Real-world agents rarely observe the full state of the world; they infer latent state from incomplete, noisy signals.
Model Components
In the standard RL objective, what does the discount factor control?
The discount factor reduces the contribution of distant future rewards.
Correct answer: The relative weighting of future rewards
What is the main difference between a fully observed MDP and a POMDP?
POMDPs introduce partial observability and require belief-state reasoning or memory.
Correct answer: In a POMDP the agent does not directly observe the full environment state.
From Policies to Plans
Agentic AI often blends reactive policies with deliberative planning. The system may use a policy for rapid action selection and a planner that searches over future trajectories to satisfy constraints or maximize long-term value.
Which formulation best captures an agent operating under partial observability?
Partial observability is classically modeled with a POMDP and belief-state inference.
Correct answer: A POMDP with belief updates