Reinforcement Learning and Sequential Decision-Making
Learning by Interaction
Sequential decision-making
Reinforcement learning studies agents that learn through interaction, balancing immediate rewards and long-term return. The canonical formalism is the Markov decision process, defined by states, actions, transitions, rewards, and discounting.
Value functions
The state-value function is , and the action-value function is . These quantities support policy evaluation and control.
RL algorithm families
| Family | Example | Strength |
|---|---|---|
| Dynamic programming | Value iteration | Exact if model is known |
| Monte Carlo methods | Policy evaluation from episodes | Simple and model-free |
| Temporal-difference learning | Q-learning | Bootstraps from partial returns |
| Policy gradients | REINFORCE, actor-critic | Directly optimizes stochastic policies |
RL is data hungry
Many RL algorithms are sample-inefficient, sensitive to reward design, and unstable under function approximation; improving these properties is a central research challenge.
Which statement best describes the exploration-exploitation trade-off?
An agent must sometimes explore to discover better actions, even if exploitation is currently safer.
Correct answer: Balance trying uncertain actions against using known good ones
Why is off-policy learning useful?
This is important when collecting new data is expensive or unsafe.
Correct answer: It allows learning about one policy while following another behavior policy, improving data reuse and flexibility.
Policy optimization styles
Value-based
- Learn action values
- Derive policy from values
- Can struggle with continuous actions
Policy-based
- Directly parameterize the policy
- Works naturally in continuous spaces
- Often higher gradient variance
Designing an RL experiment
-
1
Step 1: Define reward, termination, and discounting.
-
2
Step 2: Choose whether a model is available.
-
3
Step 3: Select an algorithm matched to action-space structure.
-
4
Step 4: Measure sample efficiency and stability.
-
5
Step 5: Evaluate robustness across seeds and environment variations.