TL;DR

If you run QA at an insurer, the jump from checking five calls a week to reviewing every interaction is not a bigger version of the same job. It means rebuilding, as explicit rules, everything a trained reviewer used to carry in their head: which SOP steps are serious enough to zero a score outright, which single record ties a policy's calls, emails and chats together, who is allowed to contest a score and where that goes, and how you stop an accent from being graded as a communication fault. Skip that rebuild and full coverage will not close your blind spots. It will simply repeat the same scoring mistakes on every call instead of on five a week. Do the rebuild first, test it against your own traffic, and coverage becomes something you can trust rather than something you merely have.

Build Your Coverage Readiness Table

Before you sign off on any move to full coverage, fill in the table below for your own operation. An answer you cannot write down yet is the gap that will show up as a scoring error later, once every call is being scored instead of a handful.

What to establish Your answer
True interaction volume per channel this month, across calls, email, chat and walk-ins
The specific SOP steps that must trigger an automatic zero rather than a weighted deduction
The one record every channel's evidence has to resolve to
Who is allowed to raise a dispute on a score, agent, supervisor, or both, and where it routes
How you will check coverage and scoring against your own live traffic before wider rollout

The sections below work through why each row matters and what it costs a team that skips one.

Count Every Channel Before You Sample

Count the true interaction volume across every channel this month, not just the one your current sample is drawn from. A sampling rate looks defensible in isolation. Set against the real mix of channels a customer actually uses to reach you, it usually is not.

A reviewer who samples a handful of calls a week builds real confidence in that number, and the confidence has nothing to do with the volume it is meant to represent. The moment you add up email, chat and walk-in traffic alongside calls, the sample shrinks to a fraction of a fraction, and most of what customers went through in a month was never looked at by anyone at all.

A Nigeria based insurance company that still runs manual review showed us exactly what that gap looks like in practice. Its QA team sampled around five calls per agent a week, across a workforce of more than a hundred customer facing staff. In the same month, the business logged between six and seven thousand emails, three to four thousand inbound calls, and up to two thousand WhatsApp messages, alongside web chat and walk-in traffic that added even more volume for the same small QA team to review manually. The team reported reviewing about five calls per agent each week. Set against that channel mix, five calls a week accounted for only a small slice of the conversations customers were actually having with the business.

If your own sample covers a fraction of a percent of true volume once every channel is counted, treat that as evidence of a blind spot, not as a working QA process.

Critical Errors Must Zero the Score

Carry forward the rule your existing scorecard almost certainly already applies: some errors are critical enough to zero the whole score, no matter how well the rest of the call went.

Weighted criteria are good at judging overall call quality, but they average. A call that handled everything else beautifully and missed one compliance step will still come out with a respectable score if that step is only worth a fifth of the total, which understates exactly the risk a regulator or a customer would care about most.

The insurer's own manual grading sheet already builds this in. Missing a customer's policy number on a call is treated as a critical error, and it zeroes the entire evaluation regardless of anything else the agent did well. When the team asked whether an automated system could replicate that logic, they were really asking whether the new process would flatten years of judgment about what matters most in a regulated business into a single average score.

A model that only adds up weighted criteria will average away the one mistake that actually matters.

One Record Ties Every Channel Together

Pick one record that every channel's evidence has to resolve to, because it is very unlikely that a single system spans your calls, email, chat and walk-ins.

Most contact centres run a voice platform for calls, a case management tool for tickets, and separate logs for chat and email. None of them, on its own, tells you what actually happened on a customer's account. Coverage across channels only holds up once you commit to one identifier that ties them together and require every review to confirm against it, rather than against whichever system happens to be open.

The insurer never adopted a sales CRM at all. Instead, its team logs every interaction, calls, emails, WhatsApp messages, web chat and walk-ins, into a case management tool against the customer's policy number. When a reviewer checks a call, they do not stop at the transcript. They confirm the policy number was captured correctly and cross check other systems, for instance confirming that a promised email actually went out, rather than trusting any one log as the full picture.

When two systems disagree about what happened on an interaction, trust the shared record over either system's own log.

Getting that shared record right before you scale coverage is the kind of criteria design worth testing on real calls rather than a slide deck, and it is what we walk teams through at https://insight7.io/book-demo/.

Match Disputes to How They Happen

Design the dispute workflow around who actually notices a disagreement first in your operation, not around a generic model where only the agent can raise a challenge.

In most self-service designs, an agent sees a low score, disagrees, and files a challenge after the fact. That is not how disagreement surfaces in every operation. At the insurer, scorecards are compiled centrally by a QA team and shared with supervisors on a weekly cadence, and supervisors pass them on to agents afterwards. Under that flow, if anyone is going to catch a wrong score early, it is the supervisor reviewing it before the agent ever sees it, not the agent contesting it days later. The team specifically asked whether a supervisor could raise a contest on an agent's behalf, because in their workflow the supervisor is the first line of defence against a bad score, not an afterthought layered on top of an agent's own challenge button.

A contest mechanism built only for the agent solves the wrong half of that problem.

Separate Accent From Communication Skill

Score accent and communication skill as two separate criteria, or full coverage will quietly grade agents on how they sound rather than what they did.

A handful of experienced reviewers can hear a strong regional accent and judge the substance of a call without thinking twice about it. Automated evaluation at full coverage does not do that unless the criteria explicitly tell it to. Every accent, cadence and turn of phrase becomes part of what a model listens for, and a criterion written loosely around tone of voice will start marking agents down for how they sound instead of for what they actually said or did.

This was a concern the insurer's team raised while discussing how the AI persona's voice matching and the evaluation criteria would handle tone of voice. Its contact centre draws on a wide range of local accents, and the team asked directly whether that variation would end up scored as a communication fault rather than left alone as simply how their agents speak. It is the same failure we see most often when operators move from spot checks to full coverage: the bias was always there in the criteria, it just never had enough volume behind it to show up as a pattern.

The moment evaluation covers every call, every agent's voice pattern becomes part of what gets scored, whether you intended it or not.

Pilot on Your Own Live Traffic

Treat a pilot run on your own live interactions as a required gate before wider rollout, not a nice to have proof of concept.The insurer put this in plain terms. It would only move forward with a pilot on its own live account and real data, not a sandbox filled with sample calls. Treat this as a standing rule based on lessons learned from past vendors who promise heaven and earth during the sales process, only to deliver far less once the contract is signed.

A demo built on someone else's clean data cannot tell a team whether its own SOPs, its own accents, and its own case management setup will actually score correctly. The pilot itself, run on their own live traffic, became the actual decision point rather than anything shown in a slide deck.

Judge coverage and scoring against your own traffic first. A vendor's demo data cannot tell you whether the criteria hold on your calls.

Questions Insurance QA Leads Ask

How many calls do we need in a pilot before we can trust the criteria? Enough to cover a full cycle of your own channel mix, not just an easy sample of clean calls. A pilot that only reviews straightforward inbound calls will look accurate and then fail once it hits collections calls, complaints, or a busy Monday. Run it across a few weeks so seasonal and channel-specific patterns actually show up before you commit.

What happens when an automated score and a supervisor's own judgment disagree? Treat every disagreement as an input to the criteria, not a one-off exception to wave away. Log the disputed score, the reason given, and the resolution, and feed that log back into how the criteria are worded. Teams that skip this step end up refighting the same disagreement every month instead of fixing the rule that caused it.

Can one evaluation criteria set cover calls, email and chat at once? No. Each channel needs its own sub-criteria under one shared framework, because the evidence looks different in each: tone and pacing on a call, response time and completeness in an email, thread continuity in chat. What has to stay constant across all of them is the shared record everything resolves to and the critical-error rules that zero a score regardless of channel.

Does full coverage mean we no longer need manual reviewers? No, their job changes rather than disappears. Once every call is scored automatically, manual reviewers shift toward auditing the edge cases the model handles badly, calibrating criteria against real disputes, and deciding which SOP steps deserve zero-score status in the first place. That judgment work is the part no rollout should skip.

Start With the Criteria, Not Coverage

If you are the one signing off on this move, resist the instinct to buy coverage first and settle the criteria later. The volume count, the shared record, the zero-score rules and the dispute path all have to exist before the first full-coverage report reaches a supervisor's desk, because a wrong criterion reviewed by a machine on every call is a mistake repeated at a scale no spreadsheet ever reached. Start by writing down the rules your best manual reviewer already applies without thinking, and only then decide how much of your traffic you are ready to cover. When you are ready to test that against your own calls rather than a demo account, book time at https://insight7.io/book-demo/.