Home › Blog › What are AI evals?
Guide · Part 1 of 2

What are AI evals, and why does agentic AI need them continuously?

Published August 20, 2026 · 10 min read

Every engineering org already runs a quality process: write the spec, prove it at a release gate, then watch uptime once it ships. That process quietly assumes the software behaves the same way every time it runs. An AI agent built on an LLM doesn't — which is why the release gate that has worked for twenty years of deterministic software doesn't bound what an agent can actually do in production, and why evaluation has to move from a pre-release checklist to a continuous layer over live traffic.

TL;DR — Traditional software is deterministic: the same input always produces the same output, so a finite test suite at the release gate can prove correctness, and production monitoring only has to confirm the already-proven code path is up and within its SLA. An LLM-based agent is probabilistic, and its effective input space — free text, conversation history, retrieved documents, other agents' tool output — isn't enumerable, so no pre-release suite can bound what it will decide to do. Evals are the practice built to close that gap: scoring an agent's real decisions against its target use case continuously, on 100% of production traffic, with a human subject-matter expert (SME) in the loop to catch the drift and hallucination automated metrics miss. This is part 1 — the case for why continual evals aren't optional for enterprise agents. Part 2 covers how to implement them.

Traditional software: a bounded, testable input space

Deterministic software has a property that makes testing meaningful in the first place: given the same code and the same state, the same input always produces the same output. A developer can enumerate the requirements, walk the known edge cases, and write a finite set of test cases that cover them. Code coverage is a real number precisely because the number of paths through a program is finite and countable — every branch, every conditional, every error handler was written by someone who had to decide what happens in that case.

That's what makes a release gate possible: the input space is bounded enough that "we tested it" is a claim you can actually stand behind.

The release gate, and what production testing actually checks

Most engineering orgs run the same two-stage quality process, whether or not they'd describe it this way:

Traditional software: proven once at the gate, then watched for uptime
Design & spec
Requirements and edge cases enumerated up front — a bounded, known input space.
→
Release gate
Unit, integration, and end-to-end tests. Deterministic pass/fail against the spec. Ships only if green.
→
Staged / canary rollout
The same already-tested code path, exposed to a slice of real traffic before full release.
→
Production
Synthetic transaction checks, uptime probes, latency/error-rate assertions against a known SLO.
Production monitoring answers "is the already-proven code path up and within SLA?" — not "was this particular output correct?" That question was already closed at the release gate.

This is worth stating precisely, because it's the assumption agentic AI breaks: production testing for deterministic software checks reliability, not correctness. Synthetic monitors, uptime probes, and SLO-based alerting (the "golden signals" — latency, traffic, errors, saturation — that underpin most SRE practice) exist to catch infrastructure and dependency failures, not to re-verify business logic that a passing test suite already proved. Sampling a fraction of traffic with a synthetic probe is a valid strategy here, because behavior is uniform: if the code path handled one well-formed request correctly, it handles the next one identically, so a canary transaction validates the whole class.

Why agentic AI breaks both assumptions

An LLM-based agent breaks the release-gate assumption and the production-testing assumption at the same time, for the same underlying reason: nobody wrote the function that decides what it does next. The model is a learned, probabilistic mapping from context to output, not a set of branches a developer authored and can enumerate.

  • The same input isn't guaranteed to produce the same output. Sampling, prompt sensitivity to small wording changes, and provider-side model updates the developer doesn't control all mean two runs of "the same" request can diverge.
  • The effective input space isn't enumerable. Free-text phrasing, multi-turn conversation history, retrieved documents, and other agents' tool output combine combinatorially. A pre-release test suite, however large, samples a sliver of that space — it cannot bound what an unknown user will type, in what order, referencing what earlier turn.
  • The known-unknowns/unknown-unknowns line moves. Traditional testing already accepts it can't cover every unknown unknown, but the known space it can cover is large relative to the whole. For an agent, the unbounded portion is the majority of the space, not the edge case.

This isn't a niche concern — it's exactly what the industry's own risk frameworks were written to address. OWASP's Top 10 for LLM Applications names overreliance on model output and prompt injection as top-tier, unresolved risk categories, not implementation bugs to be patched away. NIST's AI Risk Management Framework builds "Measure" in as a continuous function alongside "Govern," "Map," and "Manage" — explicitly because a one-time assessment doesn't hold for a system whose behavior isn't fixed by its code.

This isn't hypothetical — it's already happened in production

In practice —
  • Air Canada, 2024: A Canadian tribunal held the airline to a bereavement-fare policy its support chatbot invented for a customer — one that didn't match the airline's actual published policy. The behavior was live in production, answering real customers, before the case surfaced it.
  • A car dealership's chatbot, December 2023: A customer prompted a US dealership's website chatbot into agreeing, in writing, to sell a vehicle for one dollar — an interaction shape no release-gate test case had anticipated.
  • DPD's delivery chatbot, January 2024: A customer got the courier's support chatbot to swear at them and call the company "the worst delivery firm in the world," in a channel built for order tracking.

None of these were caught by pre-release testing. Each surfaced because a real user, in production, found an input nobody had enumerated at design time — and each ran for a real customer before anyone at the company knew it had happened.

What "evals" means for an agentic system

Evals are the practice built for exactly this gap: instead of a finite test suite that proves correctness once, evals continuously score an agent's real decisions — the actual tool calls, actions, and responses it produced for real users — against the behavior it's supposed to have. The release gate doesn't disappear; it stops being sufficient on its own, because it can only ever cover the fraction of the space anyone thought to write a test case for.

Agentic AI: the gate can't bound it, so evaluation has to run continuously in production
Unbounded real-world input
Free text, conversation history, retrieved documents, other agents' tool output. Not enumerable.
→
Agent reasons & acts
A learned, probabilistic function decides the next message or tool call. The same input can yield a different action.
→
Continual eval layer
Scores 100% of production traffic: task success, groundedness / hallucination, policy and action-boundary compliance.
→
SME / human-in-the-loop review
Flagged, low-confidence, or high-stakes trajectories get human judgment against the original target use case.
The SME's verdicts feed back into eval criteria and guardrails, which re-score the next batch of live traffic — a closed loop, not a one-time gate.

The "100%" matters and isn't a stylistic choice. Synthetic monitoring can sample because deterministic behavior is uniform across well-formed requests. An agent's behavior isn't uniform — the highest-risk interaction is often the one furthest from anything a sample would happen to catch, precisely because it's the one the design phase didn't anticipate. Sampling production traffic for an agent means the tail you most need visibility into is the tail you're least likely to see.

Determining true intent in multi-turn enterprise conversations

Enterprise agents rarely resolve a request in one turn. A support agent might spend six exchanges narrowing down an account, confirming a policy exception, and only then issuing a refund or updating a record. Scoring each message in isolation misses the thing that actually matters: whether the trajectory — the accumulated context, corrections, and clarifications across the whole session — actually supports the action the agent eventually took.

An action can be well-formed and still be wrong given the full conversation: a refund issued because the agent inferred authorization from an ambiguous earlier turn, a record updated because a clarifying question was answered in a way the agent over-generalized from. Determining the user's true intent — not just what the last message said, but what the whole conversation actually established — is a session-level judgment, and it's exactly the kind of semantic, context-dependent call that a per-message automated check is structurally not positioned to make on its own.

The human in the loop: why an SME still has to watch

Automated scoring — LLM-as-judge grading, groundedness/factual-consistency checks against retrieved sources, rule-based policy checks — approximates quality at a scale no human review queue could match. But approximation is the operative word: a fluent, confident, well-formatted response can still be a hallucination, and an automated judge trained on general criteria doesn't necessarily know that this enterprise's target use case draws the line somewhere specific to their business.

A subject-matter expert closes that gap: reviewing flagged, low-confidence, or high-stakes trajectories, confirming or overriding the automated verdict, and feeding the correction back into the eval criteria and guardrails so the same miss doesn't recur. It's the same verify-and-improve loop SRE teams already run over incident postmortems — applied to agent decisions instead of infrastructure incidents. And it doesn't end at go-live: user phrasing shifts, upstream model providers change behavior, new integrations get added, and the target use case the agent was built for can quietly drift out from under it if nobody keeps checking that production behavior still matches original intent.

Two models of quality, side by side

DimensionTraditional deterministic softwareAgentic AI (LLM-based)
Input spaceFinite, enumerable — requirements and edge cases can be listedEffectively unbounded — language × history × retrieved context × tool output
Same input, same output?Yes, by constructionNot guaranteed — sampling, context sensitivity, upstream model changes
What proves correctnessA passing test suite against a written specNo suite can enumerate the space; correctness is probabilistic and contextual
Pre-release testing goalProve every known code path behaves as specifiedCatch gross failures on a sample; cannot bound unknown inputs
Production testing goalConfirm the proven code path is up and within SLADetermine whether each real decision was actually correct
Sampling acceptable?Yes — behavior is uniform, so a probe validates the classNo — the highest-risk inputs are the ones a sample is likely to miss
Who signs off, and whenQA/eng, once, at the release gateAn SME, continuously, against real production trajectories

What's next

None of this is an argument against release gates or synthetic production monitoring for the deterministic parts of an agentic system — the API layer, the tool integrations, the surrounding application still benefit from both, and still should have them. The argument is narrower and specific to the model-driven decision itself: for the part of the system where an LLM decides what to do, the release gate can't be the last checkpoint, because it was never built to bound a space this large. Continual evaluation of live production traffic, with an SME in the loop, is what closes that gap.

Part 2: how to actually implement continual evals

This post makes the case for why enterprise AI agents need evaluation to run continuously against real production traffic, not just at a pre-release gate. Part 2 covers how to build that loop in practice — including how Parapet enables it — and will publish as a separate follow-up post once this one is live.