Beyond Pass@1: Building a Production Reliability Harness for AI Agents

Pass/fail accuracy is not enough for production agents. Build a reliability harness that measures consistency, robustness, predictability, and bounded failure harm.

Mapki

Mapki

Oct 7, 2026•8 min read
Share:
Beyond Pass@1: Building a Production Reliability Harness for AI Agents

Beyond Pass@1: Building a Production Reliability Harness for AI Agents

A tool-using agent can succeed once and still be a production liability. The same request may produce a different decision on the next run, a small wording change may derail the plan, or a rate limit may turn a recoverable API failure into a destructive partial action.

The engineering mistake is treating task success as the whole quality signal. Production systems need a broader contract: repeated behavior should be understandable, perturbations should not cause unexplained collapse, uncertainty should be visible, and failures should be bounded.

This article shows how to turn those requirements into a practical reliability harness for agents that call APIs, browse systems, or orchestrate multi-step workflows.

The gap between capability and reliability

The 2026 State of Agent Engineering report from LangChain surveyed more than 1,300 professionals and found that 57% already had agents in production. Yet quality remained the leading barrier, cited by 32% of respondents. The same report found that observability adoption was much higher than evaluation adoption: nearly 89% had some observability, while 52.4% ran offline evaluations and 37.3% ran online evaluations.

  • That pattern is familiar: teams can see traces, but they do not always have a durable definition of what a good run looks like.

  • A pass rate answers one question: Did this test episode end in an acceptable state? It does not answer:

  • Consistency: Does the same task produce stable outcomes across repeated runs?

  • Robustness: What happens when inputs, tool responses, or external state vary slightly?

  • Predictability: Does the agent know when it is uncertain and escalate appropriately?

  • Safety: If the agent fails, is the severity bounded?

  • Resource behavior: Are token, latency, and tool-call costs predictable enough to budget?

What recent research changes in the test plan

ReliabilityBench proposes a production-like evaluation surface instead of a single score. It measures repeated execution with pass^k, semantic task perturbations, and controlled tool/API failures. Across 1,280 episodes, the paper reports that small semantic perturbations reduced success from 96.9% at zero perturbation to 88.1% at a moderate perturbation level. Rate limiting was the most damaging injected fault in its ablations.

A broader 2026 study, Towards a Science of AI Agent Reliability, decomposes reliability into four dimensions: consistency, robustness, predictability, and safety. Its key warning is that capability improvements do not automatically produce proportional reliability improvements. Open-ended environments are especially difficult because the agent must navigate changing external state rather than a tightly controlled simulator.

  • The practical conclusion is simple: test the operating envelope, not just the happy path.

A reference architecture for the harness

A useful harness has five layers:

  • Scenario registry: Versioned tasks, policies, fixtures, expected end states, and risk levels.
  • Runner: Executes the agent with a fixed seed or controlled randomness, records every tool call, and enforces budgets.
  • Fault injector: Applies timeouts, rate limits, partial responses, schema drift, and delayed results to test recovery.
  • Evaluator: Checks end-state equivalence, policy compliance, trajectory diagnostics, resource usage, and harm severity.
  • Replay store: Saves traces and failing episodes as regression tests for CI and online monitoring.

The harness should evaluate state transitions and side effects, not merely compare final prose. Two agents may use different plans yet reach the same safe end state. Conversely, a fluent final answer can conceal an unauthorized API call.

from dataclasses import dataclass
from statistics import mean, pstdev

@dataclass
class Episode:
success: bool
safe: bool
tool_calls: int
latency_ms: int
cost_usd: float
confidence: float
expected_confidence: bool


def coefficient_of_variation(values):
avg = mean(values)
return pstdev(values) / avg if avg else 0.0


def summarize(episodes):
return {
"pass_rate": mean(int(e.success) for e in episodes),
"safety_rate": mean(int(e.safe) for e in episodes),
"tool_call_cv": coefficient_of_variation([e.tool_calls for e in episodes]),
"latency_p95_input": sorted(e.latency_ms for e in episodes)[max(0, int(.95 * len(episodes)) - 1)],
"cost_cv": coefficient_of_variation([e.cost_usd for e in episodes]),
"confidence_calibration": mean(int(e.expected_confidence) for e in episodes),
}

In a real implementation, replace the simple percentile calculation with a production statistics library, persist the full trace, and separate deterministic checks from model-based judges.

1. Measure consistency with repeated runs

Run each scenario at least five times under nominal conditions. Record three levels of consistency:

  1. Outcome consistency: Does the agent reach an equivalent business state?
  2. Trajectory consistency: Does it select roughly the same classes of tools and recovery steps?
  3. Resource consistency: Do latency, tokens, and tool calls stay within a predictable range?

Trajectory diversity is not automatically bad. Exploration can help in open-ended research. Treat trajectory consistency as a diagnostic unless the workflow has strict audit or rollback requirements.

For high-risk workflows, calculate pass^k: the probability that all k repeated attempts succeed. A 95% single-run success rate becomes only 77% across five independent repetitions. That is the difference between a demo metric and an operational SLO.

2. Test robustness with semantic perturbations

Generate controlled variations that preserve intent:

  • Rephrase the request without changing the desired outcome.
  • Change irrelevant ordering, formatting, or punctuation.
  • Add harmless context noise.
  • Vary identifiers, dates, and pagination boundaries.
  • Return equivalent API data with different field ordering.

The evaluator should compare end-state equivalence, not string similarity. For example, a scheduling agent can produce different confirmation prose while still creating the same approved event with the same participants and time.

Track a perturbation curve rather than one aggregate number:

scenario: refund_request
nominal_success: 0.97
perturbation_levels:
- epsilon: 0.00
success: 0.97
- epsilon: 0.05
success: 0.95
- epsilon: 0.10
success: 0.92
- epsilon: 0.20
success: 0.88
acceptance:
max_drop_at_0.10: 0.08
must_escalate_on_policy_ambiguity: true

A steep drop indicates brittle context handling. Fix the contract or orchestration before increasing model size.

3. Inject failures like a chaos engineer

Agent systems are distributed systems with a language model in the control loop. Test them accordingly:

  • Timeouts: Does the agent retry safely, or duplicate an action?
  • Rate limits: Does it back off, switch providers, or escalate?
  • Partial responses: Does it validate completeness before acting?
  • Schema drift: Does it reject unknown or missing fields safely?
  • Stale state: Does it re-read authoritative state before committing?
  • Tool authorization failure: Does it stop without leaking credentials or bypassing policy?

ReliabilityBench specifically highlights rate limiting as a high-impact failure mode. In practice, the most important assertion is often not “the agent eventually succeeds,” but “the agent never performs the irreversible step twice.”

Use idempotency keys, transactional outboxes, compensating actions, and explicit commit stages for tools with side effects.

4. Make predictability and uncertainty observable

An agent should expose signals that allow the system to distinguish confidence from guesswork. Useful signals include:

  • Policy confidence: Is the request clearly inside the allowed policy?
  • Tool confidence: Is the selected tool and argument schema unambiguous?
  • State confidence: Has the agent observed fresh authoritative state?
  • Escalation reason: Which uncertainty or risk threshold triggered handoff?

Do not blindly trust a self-reported probability. Calibrate confidence against held-out scenarios and measure whether low-confidence episodes actually fail more often. When confidence is poorly calibrated, use deterministic rules to force review for high-impact actions.

5. Bound harm instead of hiding failure

A safe harness separates failure frequency from failure severity. A wrong draft email is not equivalent to an unauthorized payment, a deleted record, or an exposed secret.

Classify violations by maximum severity per episode:

  • Low: recoverable formatting or routing error.
  • Medium: incorrect but reversible state change.
  • High: irreversible action, policy violation, privacy breach, or credential exposure.

Then gate deployment on both dimensions: a low violation rate and a hard ceiling on high-severity outcomes. Red-team scenarios should verify that the agent stops, asks for approval, or routes to a human when the risk boundary is crossed.

Observability is not evaluation

The GitHub agent-observability topic shows the ecosystem moving toward full-lifecycle tools: Coze Loop for development and evaluation, Logfire and Laminar for traces, and projects such as Tracely for turning production failures into replayable regression tests.

These tools answer what happened. Your evaluation suite must also answer whether what happened was acceptable.

A mature loop looks like this:

  1. Capture traces with tool arguments, state snapshots, latency, cost, and policy decisions.

2. Cluster failures by root cause rather than counting every run separately.

3. Freeze representative failures into hermetic regression cases.

4. Replay them in CI after every prompt, model, tool, or policy change.

5. Sample production traffic for online evaluations and human review.

6. Promote only when reliability and safety gates remain green.

References & Community Insights

Final checklist

Before calling an agent production-ready, ask:

1. Have we measured pass^k, not only single-run accuracy?

2. Do equivalent inputs and API responses preserve the same safe end state?

3. Have we injected rate limits, timeouts, schema drift, and stale state?

4. Can we explain and calibrate uncertainty?

5. Are high-severity failures impossible, blocked, or reliably escalated?

6. Does every important failure become a replayable regression test?

Production reliability is not a property the model gives you for free. It is an engineered system boundary around the model: observable, testable, replayable, and explicit about what happens when the world is messy.

Want to implement this in your business?

Mapki designs bespoke AI agents, custom workflow automations, and tool-agnostic integrations tailored specifically to your existing ERP, CRM, and databases.

Tags:AI AgentsAgent ReliabilityLLM EvaluationObservabilityProduction EngineeringTool Use