Insights

AI call scoring in contact centers: what your QA team does next

There’s an understandable question doing the rounds in contact centers: if the machine scores every conversation, what does my QA team spend its time doing?

The role is changing. Fast. The teams that lean into that change are finding they can do things that were never possible when evaluators were buried in spreadsheets scoring three calls per agent per month.

In Session 2 of evaluagent’s AutoQA Summer Sessions, Pete Dunn (Transformation Manager) and Chris Mounce (Product Enablement Manager) tackled exactly this: what the QA function actually looks like in an AI-enhanced world — and why it’s a better job than it’s ever been.

Here’s what they covered.

The old model wasn’t working — and most teams know it

Manual QA, for most contact centers, has meant spreadsheets, random sampling, and an endless chase for volume. Pete Dunn started in QA 120 years ago. His description of the job then: head down, chasing sample, never quite getting enough done to prove anything meaningful.

Three calls per agent per month. Month-old feedback. Evaluators so ground down by the end of the week that the quality of their scoring — and their coaching conversations — dropped off. And the agents on the receiving end? They knew the QA team as the “business police.” Already on the defensive before the coaching even started.

That’s not a QA function that’s adding a ton of value.

AutoQA can fundamentally change what that process looks like inside your contact center.

Deep dive on the calls that actually matter

When AI scores 100% of conversations, your evaluators are no longer randomly picking three calls and hoping they’re representative. The system surfaces the high-risk, high-value interactions — and sends those to your QA team.

Escalated complaints, repeat callers, conversations flagged for vulnerability. All the interactions that could move the needle on performance or expose significant risk.

Before AutoQA, you had to choose: either maintain your BAU sampling, or stop doing it and go deep on high-risk work. You couldn’t do both. Now you can. The AI handles the volume, and your evaluators handle the stuff that genuinely needs a human eye. That’s a meaningful shift — from simply trying to cover ground, to doing more meaningful work.

QA teams will oversee bots

When web chat arrived, contact centers had to figure out how to QA multiple simultaneous interactions. Now AI agents are having thousands of conversations at once, and the stakes are considerably higher.

A human agent makes a mistake on one call and moves on. A bot running the wrong instruction makes that same mistake across every conversation it handles. The horror stories are already out there — AI agents selling trucks for a dollar, booking passengers onto flights that don’t exist. These are brand-damaging in ways that are hard to recover from.

AutoQA gives you the ability to monitor those bot conversations in near real-time. You find out when something goes wrong immediately — not when you’re tagged in a LinkedIn post about your bot going rogue.

The principle here: if you’re investing in AI agents at scale, you need oversight technology that scales with it – and a team to help improve it. This is one of the biggest areas your QA team can grow into.

Customer journey analysis is now on the table

QA has traditionally looked at individual interactions in isolation: opening, discovery, resolution, close. But that’s not how customers experience your contact center. They call back. They switch channels. A problem that started on Monday’s chat might still be unresolved by Thursday’s call.

With Consumer Duty in force and the FCA sharpening its focus — they noted last year that 49% of customers could be considered vulnerable — regulators want to see how you’re treating customers across their entire journey, not just in a single scored snapshot. Automated journey analysis lets you do that at scale: identify where customers are hitting friction, flag how vulnerable customers are being handled, and present a treatment plan alongside the evidence.

As Pete put it: being able to show the FCA where things went wrong and what you did about it is almost always the better path. Closing the door and hoping it never comes out is not a strategy.

The new QA skill set: prompt engineering, calibration, and managing the AI

This is where it gets genuinely interesting for QA professionals. The skills that make a great evaluator — deep subject matter knowledge, attention to nuance, the ability to calibrate across a team — get redirected.

Writing effective prompts for a large language model is not the same as writing guidelines for a human evaluator. You can assume a human QA understands what empathy looks like in context. An LLM follows instructions literally. That means your prompts need to be precise.

QA teams are increasingly being asked to own this. To become prompt engineers, using their operational expertise to shape how the AI behaves, what it looks for, and how it frames its output. They won’t be amazing at it overnight, but they are among the most well-positioned people in the business to learn it. Tools like Line Item Builder help them develop the skill quickly.

Calibration still matters too. Making sure the whole team aligns on what the AI is producing, catching where it drifts from the intent, building confidence in the results. That’s where QA teams start to add more and more value.

A note on headcount

One question came up live: are organizations reducing QA team sizes once AutoQA is in place?

Pete’s answer was direct: Teams that have cut headcount too early have paid for it. This technology is about redeployment, not reduction. Some organizations have actually grown their QA function — bringing in business analysts and data capability — because the value being generated makes it easy to justify.

The QA team isn’t going anywhere. But the job is changing — and the teams that embrace that change stop being a cost the business has to carry and start being the function that tells leadership what’s actually happening in its customer conversations.

This article is based on Session 2 of the AutoQA Summer Sessions. Watch the recording →

Ready to explore AutoQA? Book a demo →

Ready when you are

Make every conversation count.

See how EvaluAgent helps the smartest contact centres turn quality into intelligence — and intelligence into outcomes.