Evals are the product
An agent without an evaluation set is a demo. How we build test sets from real cases, choose between pass@k and pass^k, and decide when an agent is ready to ship.
Every agent project eventually hits the same meeting. Someone says "it looks good," someone else says "it got one wrong yesterday," and nobody can say whether the new version is better than the old one.
The way out is an evaluation set, and Anthropic's engineering team has just published one of the most practical guides to building one: Demystifying evals for AI agents. Its advice lines up closely with how we work, so here is our version, with their best lines.
Start small, start from failures
"20-50 simple tasks drawn from real failures is a great start."
We agree. A first evaluation set does not need thousands of cases. It needs the cases that matter: the common path, the known edge cases, and every failure anyone has reported. Each case gets an expected outcome that a reviewer has signed off.
Choose the right reliability metric
The guide recommends pass@k when one success out of k attempts is enough (for example, generating options for a person to choose from), and pass^k for customer-facing agents that need to succeed every time.
The distinction comes from tau-bench, a 2024 benchmark of agents working with tools and simulated users. Its authors found even state-of-the-art function-calling agents succeeded on fewer than 50% of tasks, and were inconsistent: pass^8 was below 25% in the retail domain. An agent that succeeds 80% of the time on one try can fail most cases when it has to succeed eight times in a row.
For back-office work (calling insurers, drafting letters, updating records), pass^k is the honest metric.
Grade outcomes, not steps
The guide recommends grading outcomes rather than exact tool-call sequences, preferring deterministic graders, and calibrating any LLM judge against human experts. In practice:
- If the outcome is a status and a reference number, check them with code.
- If the outcome is a letter, use a rubric and an LLM judge, and measure how often the judge agrees with a person before trusting it.
- Do not fail an agent for taking a different, valid route.
Make the eval the release gate
The evaluation set is only valuable if it decides things. Ours gates every change: a new prompt, a new model version, a new tool. If the score drops, the change does not ship. That single rule removes most of the "it looks good" meetings.
Sources
- Demystifying evals for AI agents, Anthropic, January 9, 2026.
- τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains, Yao et al., arXiv, June 17, 2024.