Safety, Reliability, and Failure Modes
What can go wrong?
Common failure modes
Agentic systems can fail by taking the wrong action, using a tool incorrectly, misunderstanding the goal, overconfidently making up answers, or entering loops. Because they can act, the impact of mistakes can be larger than in a passive chatbot.
Risk terms
Action amplifies error
If an incorrect answer is only displayed, the damage may be small. If the same error triggers an external action, the damage can be much larger.
Mitigation strategies
To reduce risk, designers use guardrails, permission limits, human approval steps, validation checks, logging, and clear stop rules. Reliability improves when every important action can be reviewed.
Which is a safety strategy for agentic AI?
Validation and approval reduce the chance of harmful mistakes.
Correct answer: Add validation checks and human approval for important actions
What is one reason logging is useful in agent systems?
Logs improve transparency and debugging.
Correct answer: It helps people review what the agent did.
Beginner takeaway
The safest agent is not the most powerful one; it is the one whose actions are appropriate, bounded, and reviewable.
SAFETY & FAILURE MODES
What goes wrong—and how to design against it
Agent keeps calling tools without progress. Cause: bad termination, vague goals.
Agent invents tool names/parameters that don't exist, causing crashes.
External data contains adversarial instructions that hijack your agent.
Uncapped loops make thousands of LLM calls. Always set max_steps limits.
Agent gets distracted by subtasks and forgets the original objective.
Agent deletes files or sends emails without human confirmation.
Common Failure Causes in Production Agents
Mitigation Strategies
| Failure Mode | Mitigation |
|---|---|
| Infinite loops | Hard max_steps limit + loop detection |
| Hallucinated tools | Force structured output; validate tool schema |
| Prompt injection | Sanitize external data; isolate system context |
| Cost runaway | Token budgets + cost alerts before each run |
| Unsafe actions | Require human confirmation for destructive ops |