Prompt workshop
What is Prompt Evaluation?
Prompt evaluation is running prompt variants against assertions and comparing the matrix to a baseline.
promptfoo is one harness. A persona with no tests cannot tell you it regressed.
Why it matters
Prompt evaluation runs variants against assertions and compares results to a baseline. It answers: did this change make the prompt better, worse, or the same?
Without evaluation, prompt changes are guesswork. A change that improves one example may break ten others.
Key takeaways
- 1Evaluation compares prompt variants against baselines.
- 2It catches regressions that manual testing misses.
- 3Every prompt change should be evaluated before deployment.
Common mistakes
- ✕Testing prompt changes on one example and assuming they generalize.
- ✕Not maintaining a baseline for comparison.
Related terms
Concept neighborhood
Terms linked from Prompt Evaluation in the glossary graph.
- Prompt Evaluation
- promptfoo
- Agent Evaluation