Insights

AI Agent testing: decide what to automate, with evidence

— The definition

AI Agent Testing is the discipline of proving, before a single customer is affected, that your agent behaves correctly across the scenarios that actually matter. It starts with deciding what to automate in the first place, and it's grounded in how your customers really talk, not just the neat paths you imagined in a workshop.

The first thing you should understand is that testing isn't just the final gate before launch. The decision about which conversations to hand an AI agent is itself part of the pre-deployment process, because you can't test the wrong use case into a good one. Getting that decision right means working with the business to identify where an agent could genuinely play a major part — and then proving it can, against the reality of what those conversations actually look like.

Why getting it wrong is so common

The failure is public and expensive. Gartner predicts that over 40% of agentic AI projects will be canceled by the end of 2027, citing escalating costs, unclear business value, and inadequate risk controls — and notes that most are “early stage experiments or proof of concepts that are mostly driven by hype and are often misapplied”. Misapplied is the key word: a large part of that failure is choosing the wrong thing to automate, then discovering the mismatch in production instead of in testing.

When something is automated without being tested against reality, the results make headlines. New York City's MyCity chatbot was found advising business owners they could break the law — telling them, in various exchanges, that they could take a cut of workers' tips, go cashless, and fire staff who complained about harassment — because it was never tested against the messy questions real business owners ask. Air Canada learned the same lesson when its chatbot invented a bereavement-fare policy that didn't exist, telling a grieving customer he could claim a discount retroactively. These are exactly the kind of thing a test grounded in the real knowledge base catches before a customer ever sees it.

The biggest barrier is visibility, not effort

It’s not like teams skimp on testing effort. They test the scenarios they can think of — building cases from their own assumptions about what customers ask and how they phrase it.

What they lack, usually because they can't see clearly into their own contact center, is how customers actually speak: their real reasons for contact, their real intent, the specific terms they use, and the awkward edge cases that never make it onto a test plan.

So the agent looks ready because it passed the scenarios you happened to write — which tells you very little about the ones you didn't.

The same blind spot undermines the automation decision itself. When the next use case is chosen in a workshop, it tends to be picked for how automatable it feels rather than whether the data says it's high-volume, low-risk, and high-impact.

Both problems share one root: no representative view of what customers actually contact you about, in their own words.

How to decide what to automate

This is where evaluagent can help. Conversation Intelligence gives you a genuinely representative sample of what's happening in your contact center today — the real reasons customers get in touch, the language and intent behind each contact, the volumes, and how your agents handle those situations now. The Discovery Report, grouped by Reason for Contact, turns that into a map you can act on:

- Volume: which intents actually make up the bulk of your contacts, so you automate where it moves the needle rather than where it's convenient.
- Risk: which intents touch money, policy, vulnerability, or compliance, so you can hold those back or wrap them in tighter guardrails.
- Impact and benchmark: which journeys influence customer experience most, and how your best human agents handle them today, so you know the bar the AI agent has to clear.

The sweet spot to automate first is the overlap: high volume, low risk, high impact. That's a decision made from evidence, with the business, rather than a suggestion in a meeting.

Testing against reality

Once you know what to automate, that same representative sample becomes the foundation for how you test it. You build test cases from the real intents and phrasings customers use — including the multi-part questions and the language variants your customers actually send, not a sanitized subset. You validate the agent's answers against what your business actually says, so a hallucination like Air Canada's is caught by comparing the agent's response to your own knowledge base rather than a generic model's best guess. And you define the bar it has to clear before launch — coverage of the real intents, the known failure modes, and the edge cases — so "ready" means ready against reality, not against the neat path.

Who this helps

Operational and CX leaders:
Choose what to automate from real volumes and real impact, not a hunch, and defend the roadmap to the board with evidence.

Conversation designers:
Build and test journeys against the real intents and phrasings customers use, so the agent is ready for the majority of demand that never makes it onto a test plan.

QA teams:
Define the standard the agent has to clear before launch, the same bar they hold humans to.

Automation and product owners:
Prioritize the roadmap by evidence, so effort lands on the use cases most likely to succeed.

The result is testing you can stand behind. Instead of hoping you automated the right thing and hoping the scenarios you imagined were the ones that mattered, you choose what to automate from evidence and prove the agent works for how your customers actually behave — fixing what's broken while it's still relatively cheap and private to fix, not after a rollback, headline, or tribunal.

Explore how evaluagent tests AI agents before launch →

Ready when you are

Make every conversation count.

See how EvaluAgent helps the smartest contact centres turn quality into intelligence — and intelligence into outcomes.