AI agent evaluation and reliability
Why a single-run score overstates an AI agent's reliability, what to measure instead (pass^k, traces, judges) and how to set a target you can defend.
Before a company lets an AI agent loose on its customers or its books, someone has to say how often it will get the job right. The usual answer is a score from one test run. That shows what the agent can do, but not whether it does it every time.
Take τ-bench, a public benchmark from Sierra, which sells customer-service agents. In its airline section, an agent has to change or cancel bookings for a simulated customer while following the airline's policy, and a script then checks whether the booking database ended up as it should. Given one attempt at each task, an agent built on OpenAI's GPT-4o succeeded 42% of the time, the equivalent of 21 of the 50 tasks. Given four attempts at the same task, its chance of succeeding on all four was 20%, the equivalent of 10 of the 50. These are averages across repeated runs of every task, not counts from a single test.
The first number measures capability, whether the agent can do a task at all, and is known as pass@1. The second measures reliability, whether it does the task every time, and is known as pass^4, or pass^k for k runs. A test that runs each task once produces only the first number, and it is easy to mistake it for the second. Both come from the leaderboard in the original τ-bench repository, which has since been replaced by τ³-bench.
| Run 1 | Run 2 | Run 3 | Run 4 | One-run test says | Four-run test says | |
|---|---|---|---|---|---|---|
| Task A | pass | pass | pass | pass | pass | pass |
| Task B | pass | fail | pass | fail | pass or fail, 50/50 | fail |
An illustration, not τ-bench data: a task the agent solves only some of the time inflates the first number and disappears from the second.
This article assumes you know what an AI agent is. If not, start with our guide to agentic AI.
Why AI agent evaluation is hard #
Errors multiply, they do not add
Take an agent that completes each step correctly 95% of the time. Over twenty steps, if the steps were independent, end-to-end success would be 0.95²⁰, or about 36%. A per-step number that sounds excellent produces an end-to-end number that does not.
The independence assumption is doing the work here and it does not hold: agent steps are correlated, because a bad retrieval at step three poisons everything downstream. Treat this calculation as an illustration under independence, not a bound or forecast. What it establishes is directional. The longer the trajectory, the wider the gap between step accuracy and task success.
The agent changes the world
An agent that books a flight has booked a flight. Re-running it is not free, and a confident closing message saying the booking succeeded is not evidence that it did, which is why the more careful benchmarks compare the end state of a database against an annotated goal state rather than scoring text, the approach τ-bench takes. The same asymmetry applies to your own harness: anything that writes to a real system needs a sandbox with resettable state.
Same input, different trajectory
Two runs of the same task take different paths: sampling temperature, tool latency, retrieval order, a transient API error that retries on one run and not the other. None of it is visible in a single-run score. All of it is visible in pass^k.
The environment drifts under you
Tools change, APIs deprecate fields, policy documents get edited, and the model is updated by a provider on a schedule you do not control. An agent that passed its evaluation in March can fail the identical evaluation in June without a line of your code changing.
What to measure and in what order #
Layer 0: Is this task verifiable at all?
Every metric list assumes something it rarely states: that the task has a checkable ground-truth end state. Many agent tasks do not.
Split your task inventory in two before choosing any metric.
- State-verifiable tasks end in a condition you can assert against: a record updated, a ticket closed and routed to the right queue, a transaction posted with the right amount, a file written with the right contents. For these, correctness is a comparison, and it is cheap, objective and repeatable.
- Judge-dependent tasks do not: a summary, a piece of advice, a drafted reply, a research brief. There is no end state to compare, so correctness has to be assessed by something: a human, a rubric, or a model.
Everything below follows from which side a task falls on. Suppose the agent handles supplier-invoice exceptions: "match this invoice to the purchase order and flag the discrepancy" is state-verifiable, while "draft the email explaining the discrepancy to the supplier" is not. Evaluating the whole workflow requires both methods.
Layer 1: Outcome
Task success against an end-state check, not a final-message check: a harness that scores the agent's own account of what it did is measuring fluency, not work.
- pass@k and pass^k measure success in at least one attempt and in every attempt, respectively. Report both, and ask vendors for the second. If a vendor cannot produce it, run k trials yourself on a small subset during evaluation. Pick k from how often the same task recurs in production. The τ-bench paper reports k up to eight.
- Cost per successful task includes human rework. If an agent fails a third of the time, its low cost per call leaves out the expense of having a person redo those tasks.
Layer 2: Trajectory
Trajectory metrics help explain why a task succeeded or failed. Track tool-call correctness, step count, loop detection, and recovery after a failed call. A twenty-step turn and a three-step turn both return one answer, so a loop is invisible in a totals dashboard. You see it only if cost and steps are attributed per step.
Layer 3: Service-level
Track latency, escalation rate, human override rate, and time to resolution. These matter commercially, but keep them outside the reliability definition above. An agent can be fast, cheap and available while being inconsistent, and conflating the two is how teams report green dashboards for a system users have stopped trusting.
What this costs, and where to economise
Running pass^k costs k times a single run, labelling trajectories costs human time, and validating a judge costs more human time. None of it is free, and an engineering lead will say so within ten minutes.
Three economies keep it affordable. Run pass^k on a stratified subset: you need enough cases per stratum to see the decay, not the whole corpus. Sample trajectory labelling rather than doing it exhaustively. And run a smaller tier in continuous integration (CI) than at the release gate: CI catches regressions fast, the release gate is thorough.
What the spend buys is worth stating, because it is easy to present this as pure cost. The programme's value is the difference between shipping at a reliability you measured and shipping at one you assumed, priced with the same four questions later in this article. If the cost of an undetected failure is low, that difference is small and you should build less of this than the article describes. If it is high, the evaluation is cheaper than one incident.
Offline and online answer different questions
Offline evaluation is reproducible, belongs in CI, and proves you did not regress on tasks you already knew about, but it cannot anticipate inputs you have not seen. Online evaluation, sampling live traffic and watching escalation and override rates, catches what the offline set does not contain, but it is not reproducible and cannot gate a release. A programme with only one of them has a predictable blind spot.
Make failures attributable #
None of the trajectory metrics above are computable without traces rather than logs, with a span per step. Each span carries the tool name, the arguments, the outcome, latency, token cost, and an identifier that lets a failed production run be pulled out and turned into a test case later. The test of adequate instrumentation: when a run fails, can you say which call failed, what it was given and what it cost, without re-running anything?
Instrument to a standard rather than to a vendor, and read the standard before building on it. OpenTelemetry's GenAI semantic conventions now live in their own repository, separate from the core conventions, with no tagged release and an unresolved schema URL in the README. Use them, keep a mapping layer between their attribute names and your dashboards, and budget for a rename.
One boundary worth stating plainly: telemetry records that 1,200 tokens came back in 850 milliseconds, not whether they were right. Evaluation sits above telemetry, not inside it, and buying an observability tool does not get you an evaluation programme.
Setting a target you can defend #
There is a genre of statistic about agent projects failing before they reach production, and you have seen a version of it. The published estimates vary by tens of percentage points, because they count different populations against incompatible definitions of "production." Some mean any live deployment, some mean deployment at scale, some mean deployment with full security sign-off. Several are vendor-published with no stated methodology, which is why none of them is quoted here.
The spread is the point. A range that wide is not a measurement, and no figure inside it can tell you whether your agent is good enough, because none of them is measuring your task, your traffic or your cost of failure.
Measure against the incumbent, not against zero
A success rate is uninterpretable on its own. A rate of 70% is excellent if the process it replaces achieves 40%, and unacceptable if it achieves 95%. So establish the baseline first, on the same task set, before the agent exists. For the invoice handler: what proportion of exceptions does the current team resolve correctly, and how often does an error reach the supplier? Once the agent is running, that measurement is much harder to take honestly, because the comparison is no longer neutral.
Value when right, cost when wrong
Four questions. They are a framing, not a formula:
- What does a success earn?
- What does a failure cost?
- What does detecting and remediating a failure cost?
- Is that cost a single number, or a long-tailed distribution?
The third is the one most often skipped, and in our experience it is usually the one that decides the answer. A 90%-reliable agent drafting marketing copy is useful: the failures cost an edit and they are obvious. A 90%-reliable agent submitting regulated filings is a different proposition, not mainly because the stakes are higher but because the failures are not obvious. The asymmetry that matters is between detected and undetected failure, which is why the detection term is not optional.
When the honest answer is "do not build an agent"
High stakes plus imperfect reliability points to an assistant with an approval step rather than an autonomous agent. That is a legitimate outcome of this analysis and it should be reachable without anyone losing face.
The related move is to narrow the task until reliability is acceptable. Narrowing works, and it has a natural stopping point: keep narrowing only as long as the task still requires non-determinism. Once a task can be specified completely, it is a workflow, and a workflow built as a workflow is more reliable, cheaper and easier to audit than the same thing built as an agent.
What public AI agent benchmarks can and cannot tell you #
Public benchmarks are useful for comparing models and useless for predicting how your agent will behave.
Four questions are worth asking of any of them. How is success verified: end state, text match, or a model judge? What is the task horizon, minutes or hours? Short-horizon results do not extrapolate to long-running work. Does it report a reliability metric at all, or only single-attempt scores? And is the task set current, with scores comparable across its versions?
The fourth is the one people skip, and the τ-bench family shows why. The original repository now directs users to τ³-bench as the current version, and its maintained repository lists more than 75 task fixes: incorrect expected actions removed, ambiguous instructions clarified, impossible constraints corrected. It also states that, on its banking_knowledge domain, results produced before and after its v1.0.1 grading update are not comparable, and that affected leaderboard submissions were re-graded.
That is a benchmark doing its job properly, and it is also a warning: a score measures a specific task set at a specific version, and both move. Quote the version or do not quote the score, including for the figures at the top of this article, which is why they are labelled.
Apply the same four questions to SWE-bench, GAIA, WebArena or whatever appears next. The questions travel, but a table of scores does not.
Building an evaluation set from your own traffic #
Source cases from production, not from imagination
Hand-written test cases encode what the team already thinks the agent should handle. Production traffic encodes what users actually send, which is different and duller. Sample across the real distribution, including the boring majority. An evaluation set made only of interesting edge cases will overstate difficulty and miss the regressions that matter. For the invoice handler, that means most of the set is ordinary quantity and price mismatches, not the three exotic cases everyone remembers.
Every incident becomes a test case
When something fails in production, the trace becomes a case in the set. This loop is what stops the same bug shipping twice, and it is the first thing to fall away when a team is busy.
How many cases, and how to know the number actually moved
Enough cases that the smallest improvement you care about sits outside your noise floor. You do not need statistics to find that floor: run the same evaluation twice against the same agent with nothing changed, and the spread between those two runs is it. Anything smaller is not a result.
That measurement takes an afternoon and it prevents a common and expensive failure: over-reading small movements. The difference between 70% and 75% on 40 cases, two cases, is noise, and teams that treat it as progress tune themselves backwards.
Checking the evaluation set against production
Check the evaluation set on a fixed schedule, with a named owner:
- Does the case mix still match the traffic distribution, or has usage moved?
- Do its scores move in the same direction as production outcomes, or have the two decoupled? Decoupling means the set has stopped measuring the thing you care about.
- When production fails a class of case the set does not contain, that is a coverage gap. The fix is a new case class, not a tweak to an existing case. The first time a supplier sends a credit note instead of an invoice, that is a new class, not a difficult example of an old one.
Freeze it, version it, hold a slice back
A frozen, versioned set is what makes a regression gate meaningful. Without it, a score change could be a change in the agent or a change in the ruler. A held-out slice stops the set quietly becoming a target to tune against.
Label the trajectory, not only the outcome
Per-step labels, covering right tool, right arguments and correct recovery, feed evaluations, fine-tuning and reward models. Outcome-only labels still support evaluation, but they train weaker reward models (Lightman et al. found per-step supervision significantly outperforms outcome supervision on maths reasoning), and they cannot tell you why anything failed.
Using LLM judges without inheriting their errors #
Use a judge only where Layer 0 says you must
Reaching for a judge because it is easy to wire up is an avoidable mistake. The invoice handler needs a judge for the supplier email and none for the purchase-order match.
Validate the judge before trusting its scores
Score a human-labelled slice and report chance-corrected agreement, not raw agreement. Raw agreement flatters a judge on any task with an unbalanced label distribution: one that always answers "pass" agrees with humans most of the time on a set where most things pass. Judge validation is itself an evaluation, with its own labelled set, owner and cadence. Budget for it, or do not use a judge.
Known failure modes
Test your judge for each of these on your own data, because their size varies by judge, by task and by model version, and a published figure for one judge does not transfer to another:
- Position bias: does the verdict change when you swap the order of the two responses?
- Verbosity bias: does the longer answer win when it should not?
- Self-preference: does the judge favour output from its own model family?
- Domain calibration: does it agree with a specialist on specialist content?
- Temperature sensitivity: does the same input scored twice produce the same verdict?
Each is a test you can run in an afternoon, and running them beats knowing what the literature reports, because what you need is the number for your judge on your data.
Judge drift is a silent regression
Pin the judge model version. When it changes, re-baseline before comparing anything to history. A judge update re-scores your past as well as your present, and a quality "improvement" that coincides with a judge upgrade is not evidence of anything.
Engineering for reliability #
Levers that raise pass^k
Narrow the task. Make tools hard to misuse: constrain the signature so invalid combinations are unrepresentable, with enums instead of free text, required fields, and server-side validation that rejects rather than guesses. Make them idempotent, so a retried posting cannot pay the same invoice twice. Constrain the action space to what the task needs. Put deterministic scaffolding around the non-deterministic step, so ordering, validation and formatting are code rather than prompt. Gate retries on a state check rather than retrying blind. Define explicit fallbacks. Put human checkpoints at the irreversible steps, not uniformly across the workflow.
Most of these narrow what the agent is allowed to do. Check whether the task still requires non-determinism before deciding to keep it as an agent.
More agents is not more reliability
Multi-agent architectures are often proposed as a reliability improvement, and the evidence is not encouraging. Cemri et al. (NeurIPS 2025) report that performance gains from multi-agent systems on popular benchmarks are often minimal, and build a taxonomy of fourteen failure modes from more than 1,600 annotated traces across seven frameworks, in three categories: specification and system design, inter-agent misalignment, and task verification. They conclude that the failures are not merely a matter of the underlying model.
Our practical reading, which is an inference rather than a finding of the study: each additional agent adds coordination surface, and coordination is where many of the failures were found. Add agents to decompose a genuinely separable problem, not to improve reliability.
Regression gates and release discipline
Run evaluations in CI against the frozen set. Use canary and shadow runs on live traffic before a full rollout. Agree rollback triggers before launch, with named thresholds, rather than negotiating them during an incident. And plan for the dependency you do not control: when the provider updates the model, your system changes without a deploy. The frozen set and the pinned judge version exist so that you can tell what moved.
Turning evaluation into evidence #
The evidence file
Evaluation output is only useful later if someone can reconstruct it. Keep, as one referenceable artefact:
- the evaluation set version and where its cases came from
- run logs with trace IDs
- the judge model and version, if one was used
- the human-review sample and its agreement statistics
- the incident-to-case trail
- who signed off, and against which target
The test: if you cannot reconstruct why you believed the agent was ready to ship, you do not have evidence. You have a memory. That holds in any jurisdiction, and is worth doing whether or not anyone asks.
What the EU AI Act requires as of October 2026
For EU-facing systems the timeline changed in 2026 and a great deal of published guidance has not caught up. Regulation (EU) 2026/1744, the Digital Omnibus on AI, is dated 8 July 2026 and entered into force on 27 July 2026. Reading its effect on the AI Act timeline, the Cloud Security Alliance's analysis sets out the changes: high-risk obligations for standalone Annex III systems move from 2 August 2026 to 2 December 2027, and for AI embedded in products already covered by EU product-safety law under Annex I to 2 August 2028, while Article 50 transparency duties (apart from a short grace period, to 2 December 2026, for marking content from generative systems already on the market), the general-purpose AI provider obligations and the Article 5 prohibitions were not deferred.
If your compliance calendar still says 2 August 2026 for high-risk obligations, it is wrong. The deferral bought calendar, not scope: the evidence an assessment will ask for, should your system turn out to be in scope, is the same evidence, and it is produced by the programme described above.
Where to start #
Three stages, defined by what you have at the end of each rather than by how long they take.
The three run in order, but they overlap: if your agent is already live, some of Stage 2 has to happen before Stage 1 can finish.
Stage 1: Decide what "good" means. Inventory the tasks and split them by verifiability. Get the incumbent baseline from whoever owns the process today, measured on the same tasks. If an agent is already in production, you need a second baseline, its current performance, and that one is harder, because nobody instrumented for it. Work through the four value-and-cost questions with someone who carries the consequences of a failure, not only with the engineering team. You are done when you have a target number you would defend to a CFO, with the baselines beside it and the detection-cost question answered.
Stage 2: Build the instrument. Pull the first cases from production traces rather than writing them. Instrument per step before you start scoring, because retrofitting traces onto an existing harness costs more than building them in. Measure your noise floor before you measure anything else. You are done when the set is drawn from real traffic and sized against that floor, traced end to end, running in CI against a frozen version, with a judge validated only where Layer 0 requires one.
Stage 3: Close the loop. Route every production incident into the set as a case. On the review cadence, check three things: whether the case mix still matches traffic, whether its scores still move with production outcomes, and whether anything in the evidence file has gone stale. You are done when that review has happened twice without anyone chasing it.
These map onto the three phases of the TriStorm approach (Strategize, Build, Transform), but the sequence stands on its own.
Frequently asked questions #
What is the difference between pass@k and pass^k?
pass@k counts a task as solved if at least one of k attempts succeeds, and pass^k only if all k succeed. pass@k measures capability and is the optimistic number common on coding leaderboards. pass^k measures reliability and is the number that predicts production behaviour. The τ-bench paper introduced it for that reason.
How many test cases do I need to evaluate an AI agent?
There is no universal number. Run the same evaluation twice against an unchanged agent: the spread between those runs is your noise floor, and your set needs to be large enough that the smallest improvement you care about sits outside it.
Can an LLM reliably evaluate another AI agent?
For tasks with no checkable end state, a model judge is often the only scalable option, but it has to be validated against human labels using chance-corrected agreement before its scores mean anything, and re-validated whenever the judge version changes. Where an end-state check is available, use that instead.
Does the EU AI Act still apply to AI agents in 2026?
Yes. Regulation (EU) 2026/1744 deferred high-risk obligations for Annex III systems to 2 December 2027 and for Annex I systems to 2 August 2028, but Article 50 transparency duties (apart from a short grace period, to 2 December 2026, for marking content from generative systems already on the market), general-purpose AI obligations and the Article 5 prohibitions were not deferred and remain in force.
Conclusion #
Until a team runs each case more than once and sets the result against its own baseline, it does not know which of the two numbers describes its agent.
If you want a second pair of eyes on an agent evaluation programme, or on whether a workflow should be an agent at all, book a conversation.


