← Insights

What production agent teams actually measure

In LangChain's survey of 1,340 practitioners, 57% have agents in production and quality is the top barrier, yet only about half run offline evals. Closing that gap is the cheapest quality win available.

  • Evaluation
  • Production

LangChain's State of Agent Engineering report is one of the better snapshots of how agents are really built. With 1,340 respondents, three numbers stand out:

  • 57.3% say they have agents in production.
  • Quality is the top barrier, named by about one third of respondents.
  • 89% have implemented some observability, but only 52% run offline evaluations on test sets and 37% run online evaluations.

In other words: most teams can see what their agents did, but only half can say whether a change made them better.

Observability is not evaluation

Tracing tells you what happened on a given run. Evaluation tells you how often the right thing happens across representative cases. You need both, but only evaluation lets you change a prompt, a tool or a model with confidence.

Anthropic's guidance on agent evals remains the best short recipe: start with 20 to 50 tasks drawn from real failures, grade outcomes rather than exact steps, prefer deterministic graders, and calibrate any LLM judge against human experts.

The loop we run

For every agent we operate, we keep a simple loop between production and the test set:

  1. Trace everything in production, with sensitive fields removed at the source.
  2. Sample and review: a fixed share of cases plus every low-confidence or escalated case goes to a reviewer.
  3. Promote failures: any confirmed failure becomes a new test case, with the expected outcome.
  4. Gate releases: every change runs the full offline set; scores below threshold block the release.
  5. Watch online metrics: completion rate, escalation rate, cost per case and reviewer agreement, by week.

The test set grows from real life, so it gets harder in exactly the places the agent is weak.

Why this is the cheapest quality win

If quality is your top barrier and you do not run offline evals, you are tuning blind. Building a first evaluation set takes days, not months. It also changes conversations with stakeholders: from "it seems better" to "it went from 84% to 91% on the cases that failed last month."

Sources

  1. State of Agent Engineering, LangChain, June 2026.
  2. Demystifying evals for AI agents, Anthropic, January 9, 2026.