Reinforcement Learning and Sequential Decision-Making
Learning by interaction
Core formulation
Reinforcement learning studies agents that learn from rewards obtained by interacting with an environment. In an MDP, dynamics are defined by states, actions, transition probabilities, rewards, and discount factor .
MDP elements
What does the discount factor control?
A smaller emphasizes short-term rewards, while a larger values long-term outcomes more strongly.
Correct answer: Preference for immediate versus future rewards
What is the Bellman equation used for?
Bellman equations underlie dynamic programming and many RL algorithms.
Correct answer: It expresses the recursive relationship between value and successor values.
Key RL methods
Value-based methods
- Learn state or action values
- Example: Q-learning
Policy-based methods
- Directly optimize policies
- Useful for continuous actions
Exploration and exploitation
An RL agent must balance exploiting known high-reward actions with exploring uncertain alternatives. Efficient exploration remains a major research challenge, especially in sparse-reward or high-dimensional settings.
Credit assignment
Learning which earlier actions caused later rewards is difficult because feedback may be delayed, noisy, and confounded by stochastic transitions.
Which method directly learns an action-value function?
Q-learning estimates the expected return of action choices in states.
Correct answer: Q-learning
Name one challenge unique to reinforcement learning.
Unlike supervised learning, the agent’s data collection affects future data and reward.
Correct answer: Exploration-exploitation tradeoff