What are AI evals, and why does agentic AI need them continuously?
Every engineering org already runs a quality process: write the spec, prove it at a release gate, then watch uptime once it ships. That process quietly assumes the software behaves the same way every time it runs. An AI agent built on an LLM doesn't — which is why the release gate that has worked for twenty years of deterministic software doesn't bound what an agent can actually do in production, and why evaluation has to move from a pre-release checklist to a continuous layer over live traffic.
- Traditional software: a bounded, testable input space
- The release gate, and what production testing actually checks
- Why agentic AI breaks both assumptions
- This isn't hypothetical — it's already happened in production
- What "evals" means for an agentic system
- Determining true intent in multi-turn enterprise conversations
- The human in the loop: why an SME still has to watch
- Two models of quality, side by side
- What's next
Traditional software: a bounded, testable input space
Deterministic software has a property that makes testing meaningful in the first place: given the same code and the same state, the same input always produces the same output. A developer can enumerate the requirements, walk the known edge cases, and write a finite set of test cases that cover them. Code coverage is a real number precisely because the number of paths through a program is finite and countable — every branch, every conditional, every error handler was written by someone who had to decide what happens in that case.
That's what makes a release gate possible: the input space is bounded enough that "we tested it" is a claim you can actually stand behind.
The release gate, and what production testing actually checks
Most engineering orgs run the same two-stage quality process, whether or not they'd describe it this way:
This is worth stating precisely, because it's the assumption agentic AI breaks: production testing for deterministic software checks reliability, not correctness. Synthetic monitors, uptime probes, and SLO-based alerting (the "golden signals" — latency, traffic, errors, saturation — that underpin most SRE practice) exist to catch infrastructure and dependency failures, not to re-verify business logic that a passing test suite already proved. Sampling a fraction of traffic with a synthetic probe is a valid strategy here, because behavior is uniform: if the code path handled one well-formed request correctly, it handles the next one identically, so a canary transaction validates the whole class.
Why agentic AI breaks both assumptions
An LLM-based agent breaks the release-gate assumption and the production-testing assumption at the same time, for the same underlying reason: nobody wrote the function that decides what it does next. The model is a learned, probabilistic mapping from context to output, not a set of branches a developer authored and can enumerate.
- The same input isn't guaranteed to produce the same output. Sampling, prompt sensitivity to small wording changes, and provider-side model updates the developer doesn't control all mean two runs of "the same" request can diverge.
- The effective input space isn't enumerable. Free-text phrasing, multi-turn conversation history, retrieved documents, and other agents' tool output combine combinatorially. A pre-release test suite, however large, samples a sliver of that space — it cannot bound what an unknown user will type, in what order, referencing what earlier turn.
- The known-unknowns/unknown-unknowns line moves. Traditional testing already accepts it can't cover every unknown unknown, but the known space it can cover is large relative to the whole. For an agent, the unbounded portion is the majority of the space, not the edge case.
This isn't a niche concern — it's exactly what the industry's own risk frameworks were written to address. OWASP's Top 10 for LLM Applications names overreliance on model output and prompt injection as top-tier, unresolved risk categories, not implementation bugs to be patched away. NIST's AI Risk Management Framework builds "Measure" in as a continuous function alongside "Govern," "Map," and "Manage" — explicitly because a one-time assessment doesn't hold for a system whose behavior isn't fixed by its code.
This isn't hypothetical — it's already happened in production
- Air Canada, 2024: A Canadian tribunal held the airline to a bereavement-fare policy its support chatbot invented for a customer — one that didn't match the airline's actual published policy. The behavior was live in production, answering real customers, before the case surfaced it.
- A car dealership's chatbot, December 2023: A customer prompted a US dealership's website chatbot into agreeing, in writing, to sell a vehicle for one dollar — an interaction shape no release-gate test case had anticipated.
- DPD's delivery chatbot, January 2024: A customer got the courier's support chatbot to swear at them and call the company "the worst delivery firm in the world," in a channel built for order tracking.
None of these were caught by pre-release testing. Each surfaced because a real user, in production, found an input nobody had enumerated at design time — and each ran for a real customer before anyone at the company knew it had happened.
What "evals" means for an agentic system
Evals are the practice built for exactly this gap: instead of a finite test suite that proves correctness once, evals continuously score an agent's real decisions — the actual tool calls, actions, and responses it produced for real users — against the behavior it's supposed to have. The release gate doesn't disappear; it stops being sufficient on its own, because it can only ever cover the fraction of the space anyone thought to write a test case for.
The "100%" matters and isn't a stylistic choice. Synthetic monitoring can sample because deterministic behavior is uniform across well-formed requests. An agent's behavior isn't uniform — the highest-risk interaction is often the one furthest from anything a sample would happen to catch, precisely because it's the one the design phase didn't anticipate. Sampling production traffic for an agent means the tail you most need visibility into is the tail you're least likely to see.
Determining true intent in multi-turn enterprise conversations
Enterprise agents rarely resolve a request in one turn. A support agent might spend six exchanges narrowing down an account, confirming a policy exception, and only then issuing a refund or updating a record. Scoring each message in isolation misses the thing that actually matters: whether the trajectory — the accumulated context, corrections, and clarifications across the whole session — actually supports the action the agent eventually took.
An action can be well-formed and still be wrong given the full conversation: a refund issued because the agent inferred authorization from an ambiguous earlier turn, a record updated because a clarifying question was answered in a way the agent over-generalized from. Determining the user's true intent — not just what the last message said, but what the whole conversation actually established — is a session-level judgment, and it's exactly the kind of semantic, context-dependent call that a per-message automated check is structurally not positioned to make on its own.
The human in the loop: why an SME still has to watch
Automated scoring — LLM-as-judge grading, groundedness/factual-consistency checks against retrieved sources, rule-based policy checks — approximates quality at a scale no human review queue could match. But approximation is the operative word: a fluent, confident, well-formatted response can still be a hallucination, and an automated judge trained on general criteria doesn't necessarily know that this enterprise's target use case draws the line somewhere specific to their business.
A subject-matter expert closes that gap: reviewing flagged, low-confidence, or high-stakes trajectories, confirming or overriding the automated verdict, and feeding the correction back into the eval criteria and guardrails so the same miss doesn't recur. It's the same verify-and-improve loop SRE teams already run over incident postmortems — applied to agent decisions instead of infrastructure incidents. And it doesn't end at go-live: user phrasing shifts, upstream model providers change behavior, new integrations get added, and the target use case the agent was built for can quietly drift out from under it if nobody keeps checking that production behavior still matches original intent.
Two models of quality, side by side
| Dimension | Traditional deterministic software | Agentic AI (LLM-based) |
|---|---|---|
| Input space | Finite, enumerable — requirements and edge cases can be listed | Effectively unbounded — language × history × retrieved context × tool output |
| Same input, same output? | Yes, by construction | Not guaranteed — sampling, context sensitivity, upstream model changes |
| What proves correctness | A passing test suite against a written spec | No suite can enumerate the space; correctness is probabilistic and contextual |
| Pre-release testing goal | Prove every known code path behaves as specified | Catch gross failures on a sample; cannot bound unknown inputs |
| Production testing goal | Confirm the proven code path is up and within SLA | Determine whether each real decision was actually correct |
| Sampling acceptable? | Yes — behavior is uniform, so a probe validates the class | No — the highest-risk inputs are the ones a sample is likely to miss |
| Who signs off, and when | QA/eng, once, at the release gate | An SME, continuously, against real production trajectories |
What's next
None of this is an argument against release gates or synthetic production monitoring for the deterministic parts of an agentic system — the API layer, the tool integrations, the surrounding application still benefit from both, and still should have them. The argument is narrower and specific to the model-driven decision itself: for the part of the system where an LLM decides what to do, the release gate can't be the last checkpoint, because it was never built to bound a space this large. Continual evaluation of live production traffic, with an SME in the loop, is what closes that gap.
Part 2: how to actually implement continual evals
This post makes the case for why enterprise AI agents need evaluation to run continuously against real production traffic, not just at a pre-release gate. Part 2 covers how to build that loop in practice — including how Parapet enables it — and will publish as a separate follow-up post once this one is live.