Reinforcement Learning and Decision Making
Learning by Acting
The RL Loop
In reinforcement learning, an agent observes a state , chooses an action , receives a reward , and transitions to a new state. The objective is to maximize long-term expected reward.
Exploration vs Exploitation
Agents must balance trying new actions to discover better strategies with choosing actions that already seem effective. This tradeoff is central to reinforcement learning.
RL Concepts
What is the agent trying to maximize in reinforcement learning?
Reinforcement learning aims to learn policies that maximize cumulative reward over time.
Correct answer: Long-term expected reward
What is the exploration-exploitation tradeoff?
Good RL agents must both learn and perform.
Correct answer: The need to balance trying new actions with using actions already known to work well.
Supervised Learning vs Reinforcement Learning
Supervised Learning
- Uses labeled examples
- Learns a direct input-output mapping
- Feedback is immediate and explicit
Reinforcement Learning
- Uses reward signals
- Learns through interaction
- Feedback may be delayed
A policy in RL is:
The policy determines how the agent acts in each state.
Correct answer: A mapping from states to actions
Why can rewards be tricky to design?
The agent optimizes the reward signal, not the human's vague intent.
Correct answer: Because a poorly designed reward can lead to unintended behavior or reward hacking.