Foundations: Planning, Search, Reinforcement Learning, and Control
Sequential decision-making
Classical view
Agentic AI inherits ideas from planning, search, and control. In the standard formalism, the agent chooses actions to maximize expected return under uncertainty.
Planning and control
Classical planners reason over symbolic states and transition models, whereas reinforcement learning learns policies from interaction. Modern agentic systems often hybridize both: explicit planning for structure, learning for flexibility.
Core concepts
A planning loop
-
1
Step 1: Formulate the goal and constraints.
-
2
Step 2: Predict outcomes of candidate actions.
-
3
Step 3: Search or optimize over action sequences.
-
4
Step 4: Execute the best action.
-
5
Step 5: Update beliefs from feedback and repeat.
Which term denotes the long-term desirability of a state under a policy?
The value function estimates expected cumulative future return.
Correct answer: Value function
Why are hybrid agent architectures common?
Hybrid systems combine symbolic structure, learned priors, and adaptive control.
Correct answer: Because explicit structure and statistical generalization complement each other
State the Bellman intuition in one phrase.
The essence is recursive decomposition of long-horizon value.
Correct answer: Optimal value equals immediate reward plus discounted future optimal value
Name one limitation of pure planning in open-world agentic systems.
Other valid answers include combinatorial explosion or poor robustness to unmodeled dynamics.
Correct answer: Model mismatch under uncertainty