Evaluation, Benchmarking, and Research Methodology
How to Evaluate an Agent
Why Evaluation Is Difficult
Agentic systems are evaluated over trajectories, not just outputs. A system can produce a good-looking final answer while taking unsafe, inefficient, or brittle intermediate steps, so evaluation must inspect the full process.
Metrics
Important metrics include task success rate, cost, latency, tool efficiency, trajectory length, recovery rate after failure, robustness under perturbation, calibration, and safety violation frequency. Multi-objective evaluation is usually necessary.
Process Metrics Matter
Do not evaluate only the final answer; evaluate decision quality, recovery behavior, and resource use.
Evaluation Design
Why is final-answer accuracy insufficient for evaluating an agent?
Agent evaluation should include process and safety, not only end results.
Correct answer: Because an agent may take unsafe or inefficient steps even if the final answer is correct
Name one metric besides task success that is important for evaluating agents.
Other valid answers include cost, safety violations, recovery rate, or tool efficiency.
Correct answer: Latency
Benchmark Pitfalls
Benchmarks can be gamed, overfit, or become outdated. Strong methodology uses hidden test sets, adversarial cases, held-out environments, and longitudinal analysis to estimate real-world generalization.
What is the purpose of an ablation study?
Ablation studies identify causal contributions by removing components one at a time.
Correct answer: To isolate the effect of a specific component