Safety, Alignment, Security, and Governance of Agentic Systems
Risk surfaces
Why agents are risky
Agentic systems can transform ambiguous goals into concrete actions, which means specification errors, prompt injection, over-permissioning, and goal misgeneralization can have real-world consequences.
Core safety themes
The safety agenda includes alignment, sandboxing, permission control, monitoring, red-teaming, interpretability, and rollback strategies. In practice, no single mechanism is sufficient.
Threat model
Research challenge
A major open problem is building agents that are both useful and corrigible under distribution shift.
What is the best description of prompt injection?
Prompt injection manipulates the agent through crafted content or tool outputs.
Correct answer: Malicious instructions that alter an agent’s behavior through its inputs
Why is sandboxing valuable for agents?
Sandboxing constrains side effects and protects external systems.
Correct answer: It limits the blast radius of mistakes
Name one alignment technique for agents.
Other valid techniques include policy constraints, reward shaping, and monitors.
Correct answer: Human approval gates
What is one reason governance matters for agentic AI?
Governance addresses accountability, auditability, and deployment constraints.
Correct answer: Autonomous actions can have outsized real-world impact