Evaluation, Benchmarks, and Scientific Methodology for Agents
How do we know an agent works?
Evaluation principles
Agent evaluation must measure more than final answer quality. Important dimensions include task success, sample efficiency, tool reliability, cost, latency, safety violations, robustness to perturbation, and recovery from failure.
Why benchmarks are hard
Static benchmarks can be gamed or overfit, while realistic environments are expensive and stochastic. The result is an ongoing tension between reproducibility and ecological validity.
Common evaluation axes
| Axis | Example metric |
|---|---|
| Effectiveness | Task completion rate |
| Efficiency | Tool calls or tokens used |
| Reliability | Success under retries |
| Safety | Policy violations |
| Generalization | Performance on held-out environments |
Methodological warning
A strong demo is not the same as a strong evaluation.
Which metric best captures agent efficiency?
Efficiency should account for the computational or interaction cost of success.
Correct answer: Tool calls or tokens used per successful task
Why is ecological validity important?
Agents should be evaluated in conditions similar to deployment.
Correct answer: Because performance should transfer to realistic settings
Give one reason a benchmark may become stale.
Other valid answers include environment drift and public leakage of solutions.
Correct answer: Agents overfit to it
What is one desirable property of agent evaluation?
Stable, repeatable evaluation is essential for scientific comparison.
Correct answer: Reproducibility