Safety & ops
What is Agent Evaluation?
Agent evaluation measures whether a system’s observable behavior meets a task spec.
Tests, traces, and human review are evaluation tools. Leaderboard theater is not required.
Why it matters
Agent evaluation measures whether a system's observable behavior meets a task spec. Without evaluation, you cannot know if your agent is improving, regressing, or just getting lucky.
Evaluation requires fixtures (known-good inputs and expected outputs), metrics (accuracy, latency, cost), and baselines (what did the previous version do?).
Key takeaways
- 1Evaluation needs fixtures, metrics, and baselines.
- 2Without evaluation, you cannot detect regressions.
- 3Leaderboard scores do not replace task-specific evaluation.
Common mistakes
- ✕Relying on vibes instead of metrics to assess agent quality.
- ✕Evaluating on too few examples to be statistically meaningful.
Related terms
Concept neighborhood
Terms linked from Agent Evaluation in the glossary graph.
- Agent Evaluation
- Agent Observability
- Testing Agent
- Eval Harness