← All writing

Your AI Agent Will Fail. The Real Question Is How Often.

Aug 08, 2026
observabilityai-agentsoftware-engineeringartificial-intelligencellmops
A branching agent system connected to a continuous observation and improvement loop

A seven-step observability cookbook for measuring agent quality, finding failures, and deciding what “good enough” means in production

AI observability is often described as a better way to find bugs. That is true, but it is the least interesting part of the story.

AI system is not something you test once and declare correct. Models change. Prompts shift. User behavior evolves. A multi-agent workflow adds more handoffs where context can be lost or a tool can be skipped.

The operational question is no longer simply, “Did it pass?” It is, “How often does it fail, in which ways, and is that rate acceptable?”

That changes observability from a debugging accessory into the feedback system for the product. Here is a practical sequence for building it.

1. Decide What Must Not Regress

Start with the behavior that matters, not the dashboard.

For a support agent, that might include correct policy answers, successful tool calls, acceptable response time, and safe escalation. For a multi-agent system, add handoff quality: Did the planner give the sub-agent enough context? Did the sub-agent return the expected structure? Did the orchestrator use the result?

Write each concern as something you can count or score:

Some checks are deterministic. A required tool either ran or it did not. Other checks need judgment. An answer can be partly relevant, mostly grounded, or technically correct but unhelpful.

That distinction matters because AI quality rarely collapses into one green check mark.

A comparison between one-time testing and a continuous test-observe-learn-improve cycle

2. Instrument the Whole Path

Once you know what matters, make the system explain itself.

Capture the user input, model and prompt version, retrieval results, tool requests and responses, agent handoffs, final output, token use, latency, and errors. Use consistent tags for environment, tenant, feature, workflow, and experiment version. Without those dimensions, a metric can tell you that quality moved without telling you where to look.

Arize Phoenix is one practical option. It accepts traces through OpenTelemetry and uses OpenInference instrumentation to represent model calls, retrieval, tool use, and custom logic. Its Python client can query spans and annotations programmatically, while the Phoenix UI lets you inspect the path of one run.

There is also a small product decision that pays for itself quickly: surface the trace ID in your application. If a chatbot gives a bad answer, let the user or support engineer copy a debug reference. A report that says “the answer was wrong” starts a search. A report with a trace ID starts an investigation.

Treat the trace like evidence. Preserve it before changing three things at once.

3. Build a Representative Baseline

Production is too late to discover that you never defined normal.

Create an evaluation dataset from your own tasks, rules, tone, tools, and failure history. Include ordinary cases, edge cases, known regressions, and high-risk examples. Then split the data along dimensions that could hide a problem, such as customer type, language, workflow, or time of year.

Seasonality deserves explicit attention. A dataset collected in March may not represent tax season, holiday traffic, annual enrollment, or a product launch. Ask two separate questions:

  1. Does this sample represent the traffic we have now? 2. Does it represent the traffic we expect across the year?

Phoenix datasets and experiments provide a concrete workflow for this. A task runs against every example in a versioned dataset, and one or more evaluators score the outputs. Repetitions are useful when the same input can produce different answers.

This is your offline benchmark. Record it before launch, along with the exact prompt, model, tools, dataset version, and evaluator versions that produced it. Otherwise, later drift will be a feeling rather than a measurement.

A modular agent architecture, error-analysis loop, metrics, and offline-to-online comparison

4. Define Thresholds, Not Perfection

A nondeterministic system will sometimes miss. The serious work is deciding how much failure each workflow can tolerate.

A creative drafting assistant might remain useful with occasional weak phrasing. A compliance workflow may require a failure rate below one in a thousand. Those are different products with different consequences, so they need different thresholds.

Use a mix of evaluators:

Human review is usually the strongest signal, but it is expensive and slow. LLM judges make wider coverage possible, but they need to be compared with human labels and checked for their own drift. The useful design is a tiered one: automate obvious checks, use an LLM where judgment is needed, and route risky or ambiguous cases to people.

Phoenix supports span annotations and experiment evaluations from code, LLMs, and humans. Those produce scores and labels. Your team still has to decide which aggregate rates are acceptable.

5. Compare Changes One at a Time

When an agent gets worse, “upgrade the model” is an attractive answer. It is often the wrong first move.

The failure may come from a prompt change, missing context, retrieval quality, a tool contract, orchestration logic, or the way one agent communicates with another. A larger model can hide one weakness while increasing cost or introducing a new behavior elsewhere.

Keep the components identifiable. Version prompts, tools, models, datasets, and evaluators separately. Phoenix’s prompt client can create prompt versions, retrieve a specific version, and attach tags such as staging or production. Its experiments can then run variants against the same examples.

That gives you a disciplined loop:

  1. Change one component. 2. Run the existing dataset. 3. Compare quality, latency, cost, and failure categories. 4. Promote only when the tradeoff is acceptable.

Offline improvement is not proof of production improvement. It is evidence strong enough to justify the next controlled step.

A production workflow that moves from design intent and benchmarks to drift detection and acceptable rates

6. Watch Production Through Sliding Windows

After release, compare live behavior with the baseline over time. A weekly average can hide a bad hour, while a tiny window can turn random variation into false alarms. Use windows that match the volume and consequence of the metric.

Track both the overall rate and the slices behind it. A 99% tool success rate can look healthy while one tool, customer group, or prompt version is failing consistently.

Drift does not only come from your code. It can enter through model updates, dependencies, hardware, upstream data, changing user behavior, or the broader supply chain. This is why traces need version tags and why production examples should flow back into the evaluation dataset.

7. Isolate, Respond, and Learn

When a metric crosses its threshold, secure the scene before investigating it.

Find the affected traces. Group them by prompt version, model, tool, agent, user segment, and error type. Reproduce the failure with the captured inputs. Then decide whether the right response is a prompt fix, tool repair, rollback, fallback path, narrower AI scope, or a human escalation.

A watcher agent can help here. It can examine another agent’s trajectory, detect a skipped step or suspicious handoff, and trigger a retry or escalation. But it should supplement deterministic guardrails, not replace them. A second probabilistic system is still probabilistic.

Design recovery before you need it:

This closes the loop. Observability finds the trace. Evaluation names the failure. Thresholds tell you when to act. Experiments show whether the change helped.

A reliable agent loop combining continuous monitoring, fallbacks, human review, and a watcher agent

The Standard Is Controlled Failure

The goal is not an agent that never makes a mistake. That promise does not survive contact with a changing model, changing data, and real users.

The better goal is a system whose failure modes are visible, measured, bounded, and recoverable. You know what one-in-a-thousand means for your workflow. You can find the trace behind the number. You can tell whether the cause was the model, prompt, tool, data, or agent handoff. And you can prove that the next version is better on the cases that matter.

That is what turns AI observability from a debugging tool into an operating discipline.

— -

Further reading:

· Phoenix Client Reference

· Experiments API

· Prompts API

· Phoenix Quick Starts

Originally published on Medium.