TL;DR

A QA scorecard, also called a call evaluation form, is the scored document a QA analyst completes after listening to a call. A usable one has eight weighted lines, a three-point scale, and a quote field beside every score. The single change that separates a form that works from one that gets argued with: score whether the agent acted on the answer, not whether they asked the question.

The form

Copy this one. It is the version we would hand a QA team starting from nothing: eight weighted lines, a three-point scale per line, and a percentage at the end. Weights sum to 100. Change the weights to match your operation before you change the lines.

# Line What it scores Weight
1 Identity and verification The agent confirmed who they are speaking to before discussing the account 10
2 Opening and stated purpose The agent named themselves, the company, and why they are calling or what they can help with 5
3 Listening and follow-up The agent followed up on what the customer actually said, not only on the script 20
4 Accuracy Every fact, figure, policy and timeline the agent stated was correct 15
5 Empathy and tone The agent acknowledged the customer’s situation in their own words 10
6 Ownership and next steps The customer left the call knowing who does what, by when 15
7 Compliance and mandatory disclosures Every disclosure required for this call type was given, in full 15
8 Resolution or correct escalation The issue was resolved, or routed to the one person who can resolve it 10

Scale. Score each line 0, 1 or 2. Zero is missed. One is partial. Two is met. Multiply by the weight, divide by 200, and report a percentage.

What we deliberately did not weight: call duration. Duration is a result of the eight lines above, not a behaviour an agent chooses. Scoring it directly teaches agents to end calls, which is not the same as teaching them to resolve calls.

One field per line that most forms omit: the quote. Beside every score, the analyst pastes the line from the transcript that produced it. A score without its quote is an opinion, and opinions get disputed as a whole rather than corrected line by line.

What the form is for

A call evaluation form exists to make two people agree. It turns one analyst’s judgement into something a supervisor can coach from, a rep can challenge, and a compliance lead can hand to an auditor. Everything on the form should earn its place against that test.

This is why the form is a governance document before it is a measurement tool. The score it produces will sit in a coaching conversation, a performance review, and sometimes a dismissal. It needs to survive being read by the person it is about.

That framing decides a lot of the design choices below. A form built only to produce a number is easy to build and hard to defend.

Score the answer, not the question

Most forms are built out of presence checks. “Did the agent ask about X?” returns true the moment the question leaves the agent’s mouth. The information inside the customer’s answer is invisible to a check like that. That answer is the entire reason the question exists.

Here is what that looks like when it breaks.

A UK healthcare provider we work with runs clinical assessments over video, and audits a sample of them by hand. Their clinical lead scored one assessor as poor and began the process of parting ways with her. The automated evaluation, running against their own form, had scored the same assessor 84%.

Both were reading the same call. On it, a patient mentioned he had moved countries at 21 and had a psychotic episode that year. The assessor asked the scheduled mental-health-history question, heard the answer, and moved on.

She never asked why he moved. She never asked what caused the episode. The checklist recorded a question asked and an answer received, which is exactly what happened.

A presence check scores the agent’s script. But a good call is made of the questions that were not on it. Every line on a form that can be satisfied by speaking rewards agents who finish the script fastest. The ones who hear something and chase it are scored identically, and sometimes lower, because chasing takes time.

The fix is not a longer form. It is a second half to each line. Line 3, listening and follow-up, gets two checks instead of one.

Check The question the analyst answers
Presence Did the agent ask the open question at all?
Depth When the answer contained something unexplained, did the agent ask about it before moving on?

And the depth check is what sets the score.

Score What it means on line 3
0 Asked nothing, or moved on from a clear opening
1 Asked, but accepted an incomplete answer
2 Asked, heard the gap, and followed it

Line 3 carries the heaviest weight on our form for this reason. It is the only line that measures what an agent does with information they did not expect.

Why this gets expensive

At that same provider, every assessor can see their own score. So the 84% was not a private measurement error. It was sitting in the system as a counter-argument on the day their lead sat down to tell that assessor she was not good enough.

A form that scores presence does not just mis-measure. It becomes the defence in the conversation you built the form to support. That is the cost of getting line 3 wrong, and it does not show up in any accuracy metric.

See how Insight7 scores follow-up depth on every call instead of a sample. Book a demo.

One form per call type

You cannot score one call against two rubrics. Decide what the call was for, then score against that purpose alone. Sales calls, service calls and collections calls need different weights, and a single blended form measures none of them properly. The hard part is not writing the second form. It is knowing, before you score, which form a given call belongs to.

A US self-storage operator ran into this directly. Their store team and their sales team needed different criteria. But every call arrived carrying the same routing label, “inbound lead”, so nothing could separate them automatically.

Their second finding was more useful than the first. The team’s real problem was closing, and their form was spending its weight on greetings and call openings. They rewrote it to weight the close.

A form allocates attention, and attention is finite. Every point you spend on opening etiquette is a point you are not spending on the behaviour that decides whether the call achieved anything. Weight the form toward the outcome you are currently failing at, and re-weight it when that changes.

Practical rule: if your telephony labels cannot distinguish the call types you want to score differently, fix the labels before you write the second form.

One scale, everywhere

Pick one unit and use it everywhere. The form’s number will be read next to KPI dashboards, coaching notes and training results, and any unit change breaks that comparison.

A US waste-management company caught this in a product review with us. Their overall call performance was reported as a percentage. A new skills view showed the same agents on a ten-point scale, 5.8 and 5.0, next to an overall figure of 60.8%.

Their customer service director rejected it on the spot. His reasoning was about the agent, not the maths. A rep who takes training and improves has to see that improvement land on the same scale their live calls are rated on.

A form is compared, not read. A single score in isolation tells nobody anything. Its meaning comes from the numbers beside it: last month, the team average, the training result. A unit change severs every one of those comparisons at once.

He asked for one further alignment, and it is the stronger version of the same point. The criteria used in training simulations should be the criteria used on live calls. Otherwise practice improves a score nobody is measured on.

Check attribution first

Confirm the call belongs to the agent named on it before you score anything. A perfectly designed form scoring the wrong agent’s call is worse than no form at all. It produces confident, specific, wrong feedback, and it does it at exactly the moment a rep is deciding whether to trust the system.

At that same waste-management company, a team lead went looking for her own team’s calls and found them filed under other people’s names. The cause was mundane. One rep had never been registered as a user, so the system could not match her calls to a person and attached them elsewhere.

It surfaced only because she went looking for one particular call. Until then the scores had been landing under the wrong names, uncontested.

Attribution failures are silent, and that is what makes them expensive. A missing agent record does not produce an error, it produces a plausible-looking scorecard under the wrong name. Reconcile your agent roster against your call log before the first evaluation, and again every time someone joins.

Build in an appeal route

Publish the route for disputing a score, and answer disputes. Reps who cannot challenge one line will dismiss the entire form. You lose the parts that were correct along with the parts that were not.

The best version of this we have seen was a customer service director introducing scoring to his own agents. He told them plainly where to raise a disagreement, and that raising one gets it reviewed and adjusted. Then he told them which parts of the feedback he personally acts on and why.

That framing did more for adoption than any accuracy improvement would have. The reps knew the form was a conversation rather than a verdict before they ever saw a score.

A QA form’s authority comes from being correctable, not from being correct. An appeal route converts a rep’s disagreement into a specific, fixable claim about one line. With no route, the same disagreement becomes a view about the whole system, and that view spreads.

How to fill one in

Work in this order. A new analyst should expect an evaluation to take noticeably longer than the call itself, and that gap closes with practice. The order matters more than the speed, because two of these steps are unsafe to skip.

  1. Confirm the call belongs to the agent named on it. Thirty seconds. Skip this and everything after it is unsafe.
  2. Decide the call type and open the matching form. If two forms could apply, the call type is not defined tightly enough.
  3. Listen once without scoring. You are looking for what the call was trying to achieve and whether it did.
  4. Listen again with the form open. Score each line 0, 1 or 2 and paste the quote that produced it. If you cannot find a quote, the score is a 1.
  5. Score line 3 last. By then you know what the customer said that the agent could have chased.
  6. Write one sentence of coaching per line scored 0. Not “improve listening”. Write the question they should have asked, in the words they should have used.
  7. Send it within 48 hours. Later than that and the agent is being coached on a call they no longer remember.

Three decisions outside the form

Where specialist judgement gets reviewed. Clinical, legal and technical calls carry judgement that no scorecard fully captures. The clinical lead above was asked to write down what separated a good assessment from a bad one. His answer: he was being asked to put eight years of university and twelve years of higher education on paper. Score the behaviours the form can see, and route the judgement call to a specialist review alongside it.

How the form learns what good looks like. The strongest calibration method is also the simplest. Take calls your team already agrees were strong and weak, with the written reasoning attached, and work backwards to the behaviours that separated them. That set is worth more than another month of rubric drafting, and it is the same method that tunes an automated scorer.

What sits outside the single call. A form scores one conversation. Repeat contacts, a missed callback and a customer’s history across a month are a different view, and they belong on a dashboard rather than on the form.

Questions QA leads ask

These are the questions that come up when a team is building or rebuilding a form, in the order they usually come up. Each answer is the short version. The reasoning behind most of them is in the sections above.

What should be on a call evaluation form for a call center?

Eight lines: identity and verification, opening and stated purpose, listening and follow-up, accuracy, empathy and tone, ownership and next steps, compliance disclosures, and resolution or escalation. Weight them to sum to 100, with the heaviest weight on listening and follow-up. Add a quote field beside every score.

How many points should a call evaluation form have?

Use a three-point scale per line: 0 missed, 1 partial, 2 met. Then convert the weighted total to a percentage. Three points is enough for analysts to agree with each other and few enough that they do not argue about the difference between a 6 and a 7. Report in whatever unit your other performance reporting already uses.

How often should you evaluate calls per agent?

Most teams manage four to six calls per agent per month by hand, which is roughly two percent of what happens on the phones. That is enough to coach an individual and not enough to make a claim about compliance. If you need to say every regulated call was checked, sampling cannot get you there.

Who should fill in the call evaluation form?

A QA analyst who did not handle the call and does not manage the agent. Supervisors scoring their own reports produce scores that drift toward their existing view of the person. If the supervisor must score, have a second analyst re-score a sample and compare. The gap between them is your calibration problem.

What is the difference between a call evaluation form and a QA scorecard?

They are the same document under two names. “Form” is more common where the output is a coaching conversation. “Scorecard” is more common where the output is a reported number. Use whichever word your team already uses, and keep one version of the document rather than two.

FAQ

Should agents see their own scores? Yes, with the quotes attached. Hidden scores get treated as arbitrary the moment one is wrong. Visible scores with evidence get argued with line by line, which is the outcome you want.

Can one form cover inbound and outbound calls? No. The purposes are different, so the weights must be different. Build the second form when you have a second call type worth scoring, not before.

How do you keep two analysts scoring the same call the same way? Calibrate on labelled examples. Take calls the team already agrees were strong and weak, with the written reasoning, and work backwards to the behaviours that separated them. This is the same method that fixes an automated scorer.

How often should the form itself change? Re-weight when the operation’s biggest problem changes, and not on a schedule. Changing weights mid-quarter breaks comparison with the quarter so far, so change them at a period boundary and say so.

Running QA for a contact centre?

If you are a QA or contact centre lead trying to score more than a two percent sample, see how Insight7 applies your own evaluation form to every call. See it on your calls in 20 minutes.