Picture the scene: messy data lands in a shiny new platform, scorecards get copied across as-is, and then someone flips the switch and wonders why the scores look wrong.
In the final session of the AutoQA Summer Series, evaluagent’s Product Enablement Manager Chris Mounce walked through the four things that distinguish a smooth migration from a painful one.
All of them are in your control before go-live.
1. Understand your data before import
It sounds obvious, but it’s surprising how many roadblocks this one can create.
Plenty of contact centres store their interactions in one platform but access them through a completely different one. When it comes to integration, they point at the tool they use every day — and that’s not where the data is. The connector does its job fine, it just connects to the wrong place.
Then there’s data quality. The integration works exactly as it should. But if the data in your source system is messy, it lands in the platform messy. That becomes unscalable, since the data is unusable. No amount of clever AI fixes a bad foundation.
Redaction is the other thing that catches people out. If conversations contain card numbers, national insurance numbers, or any personal data, configure redaction before the import. Retrofitting redaction after sensitive data has already landed is fixable — but it’s work you didn’t need to do, and it adds time to a project that might already be feeling like a lot.
One more thing: this part of the migration doesn’t have to be a QA job. Usually, it’s an IT and technical job. Bring those people in early and let them work directly with your TSMs. You don’t need to learn what an API is — just get the right people involved.
The fix: Identify your true source system first. Audit data quality before you integrate. Configure redaction before a single conversation lands.
2. Map your agents, metadata, and bots
Data landing in the platform and data being usable in the platform are two different things. This is where migrations can go wrong without anyone realising it immediately.
Agent mapping. Your telephony platform and your evaluagent login both identify agents — but they may do it differently. Those conversations could come in unmatched. Nobody knows whose call it is. Get clear on how agents are identified in your source platform before you create your users, and it solves itself.
Metadata. Your source system might have 150 metadata fields. The vast majority of those are noise when it comes to QA. If you bring them all across, you lose visibility of the ones that actually matter — things like channel, business category, or VIP status, which are exactly what your scorecards and auto-scoring rules hang off. Document the fields you need. Leave the rest behind.
Multi-leg calls. If a conversation passes through three agents, you have a genuine decision to make: do you QA the whole interaction, just the last agent, or every agent individually? There’s no universally right answer. But it’s a decision that affects how you structure and attribute scoring, so make it at the start, not six months in.
Bots. If you don’t explicitly identify your bot, it could be treated as a human agent — sitting in someone’s team average, distorting KPI reports, or getting scored against a human scorecard. With AI-handled conversations increasing, this matters more than ever.
The fix: Map your agent IDs before importing. Choose only the metadata fields that serve your QA programme. Decide how to handle multi-leg calls. Identify your bots explicitly.
3. Rewrite your scorecards — don’t copy them
This is the assumption that causes the most pain: that you can lift your existing scorecard and drop it straight into AutoQA.
You can’t.
That’s because your current scorecard was written for a human reviewer. A human brings context, judgment, and instinct. They know what “Was the agent empathetic?” means in practice. Meanwhile, an AI does exactly what it’s told, very literally, thousands of times.
If the instruction is vague, the scoring will be inconsistent. If the explanation is complex — the “if-this-then-that” rules that have been living in a spreadsheet for years — it’ll break the moment it hits a system bound by real logic.
Before you migrate your scorecard, audit it. If it’s got 100+ line items, ask honestly: are all of these intentional, or have they just accumulated? Strip out what isn’t earning its place.
Then rewrite your criteria as prompts. This is a different skill.
Compliance questions translate well — did they say it or not? Subjective criteria typically need more work. What does empathy actually look like for your business? What specific behaviours are you measuring? Define it. Test it. Iterate.
evaluagent’s Line Item Builder does a lot of this heavy lifting here — it’ll write the prompt, write the guidelines, and help you iterate. But the thinking has to come from you.
And don’t forget your knowledge base’s place in QA. Your policy documents, process guides, and knowledge articles are almost certainly sitting there unused (by QA, at least). Ground your AutoQA evaluations against them and you’re scoring against what your business actually says good looks like — not a generic idea of it.
The fix: Audit your scorecard before you migrate it. Rewrite criteria as prompts. Define subjective measures precisely. Feed in your knowledge base.
4. Run in parallel — don’t flip the switch cold
This is the one that has the biggest human impact.
Imagine going live with AutoQA, and your reviewers start seeing scores they don’t recognise. They don’t agree with them. They can’t account for them. If they decide the tool is wrong, or untrustworthy, it’s very hard to come back from — you’re arguing instead of improving, and you’ve lost credibility in your QA program.
People can stand behind a number or decision they don’t understand.
The fix is reasonably simple: run AutoQA alongside your existing manual process. Don’t switch anything off. Let AutoQA run in the background, let humans check its work while trust is being built, and use that period to refine your prompts and guidelines. When the scores align — when you can prove it — then you scale.
One evaluagent customer did exactly this. A few weeks of running in parallel, iterating their scorecard, validating results. When they went live, they went from 900 manual evaluations a month to 6,000 — with 95% coverage. Read the full case study.
Don’t forget to bring your agents into this process too. Give them a way to flag where scores don’t feel right. Their feedback directly informs prompt improvements — and if they’ve been part of building the programme, they become your champions when it rolls out. Agent champions are gold. They spread the good word in ways no internal comms ever will.
The fix: Keep manual QA running. Let AutoQA prove itself alongside it. Involve agents from the start. Go live with evidence, not optimism.
The bottom line
A well-managed AutoQA migration is a human project with some technical steps in the middle. It doesn’t have to be a total headache.
Get your data clean. Map everything before you import it. Rewrite your scorecards properly. Run it in parallel until you trust it.
Skipping any of these steps will cost you more time on the other side than doing them would have cost you upfront.
And you don’t have to figure it out alone — evaluagent’s dedicated Technical Success Managers and enablement team are there to carry the technical weight and make sure you’re getting value from day one.
This article is based on Session 4 of the AutoQA Summer Sessions. Watch the recording →
Ready to explore AutoQA? Book a demo →