Build representative cases.
Agree success criteria and realistic test data. Include edge cases, denied access, ambiguity and adversarial inputs.
Measure what the system retrieved, which tools it used and what actually happened. Repeat the important cases when anything changes.
Explore the capabilityA fluent answer can hide a failed task. Evaluation should include retrieval, permissions, actions, cost and failure handling as well as response quality.
Agree success criteria and realistic test data. Include edge cases, denied access, ambiguity and adversarial inputs.
Use automated checks and human review where each is useful. Keep a record of model, prompt, data and configuration versions.
Set release gates and review responsibilities. Monitor production behaviour and investigate regressions rather than trusting a one-off score.
Change a model, rerun the same retrieval and tool-use tests, then review the differences before release.
No single score proves safety or suitability. Make the test coverage, limitations and residual risks visible.
Many checks can be. Human judgement remains useful for ambiguous quality questions, high-impact outcomes and reviewing the tests themselves.
No. Technical controls, named responsibilities, documented decisions and ongoing review need to work together.
Technical reference: Anthropic: evaluating AI agents (opens in a new tab)