← Back to CoursesAgentic AI: PhD Level

Neuroanatomy Explorer

Drag to rotate · scroll to zoom · click regions to explore

View
Loading 3D model…

Click a region
to explore it

Memory Deck

Flip each card and rate whether you knew it. Your score is saved.

Term
Definition

Deck complete — score saved.

Match the Pairs

Match each term to its definition. Finish the board to earn your score.

All matched — score saved.

The Agent Loop in 3D

Watch a thought travel through Perceive → Plan → Act → Observe. Drag to rotate, scroll to zoom, click a node.

Click a node to read its definition.

Safety, Alignment, Security, and Governance of Agentic Systems

Manual: General · Subject: Agentic AI

Investigate the risks posed by autonomous systems and the methods used to align, constrain, and govern them.

Risk surfaces

Why agents are risky

Agentic systems can transform ambiguous goals into concrete actions, which means specification errors, prompt injection, over-permissioning, and goal misgeneralization can have real-world consequences.

Core safety themes

The safety agenda includes alignment, sandboxing, permission control, monitoring, red-teaming, interpretability, and rollback strategies. In practice, no single mechanism is sufficient.

Threat model

Prompt injection
Malicious instructions hidden in retrieved content or webpages
Tool abuse
Misuse of APIs or side effects
Specification gaming
Optimizing the metric instead of the intent
Data exfiltration
Unauthorized leakage of sensitive information
Runaway autonomy
Unbounded or poorly constrained action loops
🔑

Research challenge

A major open problem is building agents that are both useful and corrigible under distribution shift.

What is the best description of prompt injection?

Why is sandboxing valuable for agents?

Name one alignment technique for agents.

What is one reason governance matters for agentic AI?