Key points
- The working answer is 1% to 3% of calls, and up to 5% in well-resourced operations. That range comes from how many analysts a contact center can afford, rather than from any statistical standard.
- Sampling 5 calls per agent per month is a common practice, and the Quality Assurance and Training Connection calls that volume statistically invalid for judging individual performance.
- The better question is how much risk stays invisible. At 3% coverage, 97 out of every 100 customer conversations pass through unscored, including every compliance breach and churn signal inside them.
- Automated QA removes the sampling decision. Insight7’s AI Call Scoring evaluates 100% of conversations against your rubric, then routes only flagged calls to a human queue, so analysts spend their hours on the calls that need judgment.
What percentage of calls do contact centers review today?
Between 1% and 3% of total call volume, with 5% at the upper end of well-resourced operations. Per-agent numbers give a clearer picture than percentages.
The Quality Assurance and Training Connection (QATC) compiled results show more than half of respondents evaluating one to four calls.
A poll by CallCentreHelper also found the most common answer among UK contact centers was more than 10 calls per agent per month, with practitioners describing very different realities. One QA lead at a 50-seat center reviews 15 calls per agent monthly as the only person doing it.
Run the math on your own floor. For example, a 50-agent support team handling 4,000 interactions a day, roughly 88,000 a month:
| Sampling approach | Calls reviewed monthly | Coverage |
|---|---|---|
| 5 calls per agent | 250 | 0.28% |
| 10 calls per agent | 500 | 0.57% |
| 20 calls per agent | 1,000 | 1.1% |
Even doubling QA headcount to hit 20 calls per agent leaves 98.9% of conversations unscored. For a CX manager at a 200-person company with no dedicated QA function, the number often sits below 0.3%.
Insight7’s own vertical research puts the same figure from the platform side.
Across scored insurance sales conversations, teams were reviewing under 2% of interactions manually. In the health and wellness study, 98% of consultations were invisible to the practices running them. Those are two separate observations that happen to land in the same place as the association survey data.
Important: Coverage below 1% means agent scorecards, performance reviews, and in some cases termination decisions rest on a fraction of a percent of someone’s actual work.
The statistical case for QA sample size and sampling methodology
QA sampling methodology has a problem that percentages hide: team-level and agent-level evaluation are different statistical questions with very different sample size requirements.
Aggregate scoring needs a few hundred calls
If the question is how the contact center performs overall, standard sampling math applies. To estimate a rate across a large population within a ±5% margin of error at 95% confidence, you need roughly 384 observations, which is why 400 shows up as a working benchmark.
For a center handling 88,000 calls a month, 400 randomly selected calls gives a defensible read on overall quality. That is achievable with a small QA team.
Agent-level scoring needs a few hundred calls per agent
Judging an individual agent means each agent becomes their own population. The same confidence math applies to each one separately, which is where sampling collapses as a strategy.
Rating agent performance on one to two calls per week is “not statistically valid” unless the results are used primarily to support coaching conversations rather than formal evaluation.
That caveat matters more than it sounds. A five-call sample where one call goes badly produces a 20% failure rate on paper. The same agent might fail 4% of the time across their real monthly volume, and the scorecard has no way to tell the difference.
What this means for how you use the score
- Coaching conversations: a small sample works. You are looking for something specific to practise, not a verdict.
- Performance reviews and compensation: a small sample does not work. The confidence interval is wider than the differences you are ranking people on.
- Compliance and regulatory audits: sampling is the wrong tool entirely, since a single missed disclosure creates exposure regardless of how the rest of the sample scored.
A QA lead at a financial services firm has to answer to auditors, and “we reviewed 3% and found nothing” is a weak position when the regulator asks about the other 97%.
Why a small QA sample size still misses what matters
Small samples fail in three specific ways, and none of them are fixed by choosing calls more carefully.
The mediocre middle disappears.
QA teams under time pressure gravitate toward flagged escalations and standout calls, which means the ordinary conversations never get reviewed.
Those calls are where slow erosion lives: the agent who stopped confirming account details, or the rep whose closing question fell away over a few weeks. Nothing about them triggers a review.
Two analysts score the same call differently
Our published calibration data shows trained reviewers disagreeing on 20 to 30 percentage points of scoring criteria when evaluating identical calls. O
n a five-call sample, that variance can swing an agent’s monthly score more than their actual performance did. Agents notice this quickly, and a scorecard they consider arbitrary stops driving behaviour.
Compliance exposure sits in the unreviewed majority
A missed disclosure on call 47 of 1,400 carries the same regulatory weight whether or not anyone listened to it.
Odun, CEO at Insight7 described the downstream effect when he said:
“Coaching happens too late. Feedback is being given after key moments have been lost and opportunities have passed. It’s impossible to detect the patterns happening in those calls at scale. How do you as an individual spot trends across hundreds of calls a week manually?”
This is where AI call scoring addresses the calibration problem specifically.

Every score anchors to the exact transcript moment that produced it, so a disputed rating gets resolved by opening the quote rather than by whoever argues harder.
How AI quality assurance software achieves 100% call coverage
AI quality assurance software scores every conversation against your rubric automatically, then routes a small subset to human reviewers based on risk. The sampling decision disappears because there is no sample.
Here is how the mechanism works in a contact center QA workflow:
- Ingestion: Recordings sync from your telephony or meeting stack. Insight7 connects to Zoom, Teams, Google Meet, Dialpad, Aircall, RingCentral, Talkdesk and Five9, so calls arrive without anyone uploading files.
- Transcription and redaction: Audio converts to searchable text across 60+ languages, with PII and PHI redaction applied before scoring for teams under HIPAA or financial services rules.
- Automated scoring: Each call runs against the rubric you define. Different call types route to different evaluation templates, so a renewal call is never marked down for missing a discovery sequence that did not apply.
- Risk-based human queue: Calls that fail compliance checks, score below threshold, or contain flagged language surface for human review. Analysts keep doing judgment work, applied to calls selected by evidence rather than by random draw.
- Calibration and override: QA leads review AI ratings, override where context justifies it, and refine the rubric. Scoring accuracy improves as the rubric tightens.
Note: Automated QA does not remove QA analysts from the process. It changes what lands on their desk from a random 3% to the specific calls that need a person.
Here’s what that full coverage looks like in practice:
TripleTen ran into the ceiling with a large customer support team trying to evaluate soft skills consistently. After integrating Zoom with Insight7 in one week, they were processing over 6,000 coaching calls a month for roughly the cost of one project manager.
The output was a standardised evaluation framework applied to every call rather than the handful a supervisor could reach.
How to choose call center QA software
Contact center quality assurance software has converged on similar feature lists, so the differences that matter are structural. Three questions separate tools that solve the coverage problem from tools that automate a sample.
Does it score 100% of calls, or still sample?
Some call center QA software automates the scoring of a sample rather than removing the sample. Ask for the coverage number on a real deployment at your call volume, not the marketing claim.
Insight7 scores every conversation that reaches the platform, with plan tiers set by analysis volume rather than by percentage.

The Business plan covers 200 call analyses monthly for teams up to three users at $299/month, and Enterprise removes the cap entirely with unlimited analyses and API access.
Can every QA score be traced to a specific transcript moment?
A score without evidence produces the same argument as manual calibration variance, moved into software. When an agent disputes a rating, someone has to be able to open the moment.

Insight7 anchors each criterion score to the transcript passage that produced it, and gives reps moment-by-moment breakdowns of their own calls. Reps accept a score they can inspect and argue with.
Does it connect QA scores to coaching or stop at a report?
A scorecard that flags weak objection handling has done half the job while someone still has to build the practice.
Insight7 links evaluations directly to coaching plans, and low scores generate AI roleplay scenarios targeting that specific gap. A rep practises against an AI buyer using the same objection, without booking a manager’s calendar.

Pro tip: build your rubric from your own calls rather than a generic template, our free Call QA Scorecard Builder generates evaluation criteria from real recordings.
Score every call with Insight7’s AI Call Scoring
If you are staffing a manual program, 1% to 3% is the realistic band and 5% is a strong result. Use those scores for coaching conversations and treat them cautiously in formal reviews, since QATC is right that a five-call sample cannot carry that weight.
When compliance exposure or agent fairness matters more than QA headcount economics, the percentage stops being the useful metric. Insight7’s AI Call Scoring evaluates every conversation against your rubric and routes only the flagged calls to your analysts, so their hours go where the evidence points.
Start by working out your current coverage. Divide the calls your team reviewed last month by your total volume and see which side of 1% you land on. Then upload a few calls and compare what automated scoring surfaces against what your sample caught.
FAQs
What percentage of calls should a call center QA team review?
Manual QA programs typically review 1% to 3% of calls, reaching 5% in well-resourced operations. For compliance monitoring, sampling any percentage leaves exposure in the unreviewed remainder, which is why regulated teams move to 100% coverage.
What is a statistically valid QA sample size?
For team-level quality estimates, roughly 400 randomly selected calls gives a ±5% margin of error at 95% confidence across a large call population. Individual agent scoring is a separate calculation, since each agent forms their own population.
Can AI review 100% of contact center calls?
Yes. AI QA software transcribes and scores every recorded conversation against a defined rubric, then routes flagged calls to human reviewers. TripleTen processes over 6,000 coaching calls monthly through Insight7 after a one-week Zoom integration, and Tri County Metals reached full coverage across more than 5,100 monthly inbound calls previously stored on local servers.
How many calls per agent should QA review each month?
Practice ranges from five to fifteen calls per agent per month depending on QA staffing. A single analyst is often responsible for 50 to 100 agents, which sets the ceiling in most programs.
Does automated QA replace QA analysts?
No. Automated scoring handles the volume, and analysts handle the calls that need judgment. The change is in which calls reach them: a risk-based queue built from compliance flags and threshold breaches, rather than a random selection that may contain nothing worth reviewing.


