TL;DR

A call calibration process is the session where a QA team and a scoring system agree, in writing, what a great call actually sounds like. The output is a short guide: named behaviors instead of vague labels, plus a fixed process for walking a disputed score back to agreement. Hand that guide to your auditors and to whoever builds your automated score, and the two stop disagreeing about the same call.

A high score, a poor assessor

One company found the gap the hard way. A clinical lead had already decided an assessor was not fit to keep working with patients. The automated score for that same assessor's calls said the opposite.

As the lead described it: "I've done several of her audits and I've deemed her like a poor assessor and we're going to part ways with her… but Insight 7 has scored her quite high."

The clinical lead scored the same call as a poor assessment and had to go back and reassess the patient it covered.

Scored high, on the way out

That contradiction was not hypothetical or theoretical. It landed in the middle of a real personnel decision. The clinical lead was already parting ways with the assessor on clinical grounds while the score kept saying she was performing well.

That is the actual cost of an uncalibrated score. It does not just misjudge a call. It sits in the system as a counter-argument on the day someone has to make a decision about a person, and the decision and the score point in opposite directions.

Where the score went wrong

The gap was not random. The automated score was reading tone and composure, the same things a patient would notice sitting across from her. Composure and clinical depth are different things, and a score built to notice the first cannot tell you anything about the second.

The assessor asked both questions and logged both answers, then moved on. The clinical lead marked that call down hard, because nobody asked why he moved or what caused the episode. Those two follow-up questions were the entire clinical picture, and a score reading for pleasantness had no way to know they were missing.

A score that reads tone cannot tell the difference between composure and competence.

Calibration is not more rules

The instinct when a score is wrong is to add another line to the form. That is not what closes this gap. "There's a bit of a calibration issue here… we need to understand how you are defining success… we can program that into the system."

A calibration session is not a new rulebook. It is the same handful of behaviors, written down specifically enough that a score can be checked against them instead of argued about in general.

See how Insight7 applies a calibrated guide like this to every call instead of a sample. Book a demo.

What a calibration guide contains

The guide has two parts. The first names the behaviors a great call contains, in enough detail that two people listening to the same call reach the same score. The second is the process for resolving the calls where they still do not.

Behavior What a great call does What a score alone misses
Follow-up on the unexplained Asks why, not just what, when an answer opens a door A pleasant, composed answer that never gets asked why
Root cause over surface tone Chases the cause behind a fact, not just the fact itself Tone reading as competence when the cause is never asked about
Written reasoning behind every score Every disputed or low score carries a specific, written reason A number with no explanation attached, which cannot be coached or appealed

The clinical lead's own framing was to look at what separates a great assessor from a weak one and check whether the score could see it, starting with whether the person was "asking great follow-up questions."

Score the follow-up, not the tone

This is not a problem unique to clinical calls. A mobility company we spoke with said plainly that they value the warmth a phone call brings. Warmth is worth keeping. It is also exactly the quality that can make a call sound complete when a real question never got asked, which is the same disparity the clinical lead described between a pleasant call and a clinically sound one.

A different industry, same failure

The same mobility company raised a version of this worry before a single call had even happened. They wanted to know whether an AI agent handling their calls would ever invent an answer that was not in the knowledge base. A fluent, confident answer and a correct one are not the same test, any more than a pleasant assessor and a clinically thorough one are.

Most of that team's calls go out to new signups who have not booked a trip yet. That is exactly the kind of call where an invented answer does the most damage, because the customer has nothing yet to check it against. A calibration guide has to name both failures, the missed follow-up and the invented answer, or it only catches one of them.

Walking a disputed score back

When an auditor and a score disagree, the disagreement should follow the same steps every time, so it produces a fix instead of a one-off argument.

  1. Write down the score and the specific line the auditor disagrees with. A disagreement about the whole call cannot be resolved. A disagreement about one line can.
  2. Find the exact moment in the call the auditor is objecting to, and write down what was actually said. In the healthcare case, that moment was the question about the move that was never followed with why.
  3. State what a great version of that moment would have sounded like, in the auditor's own words, not the vendor's.
  4. Add that behavior to the guide as a named line, not a general note. A note that just says to ask more follow-up questions is not a line. A line that says to ask why, when someone names a life change without giving a reason, is.
  5. Write the reasoning down in full, the same way a good audit already writes detailed feedback for a poor score.
  6. Send the updated guide to whoever scores the next call, human or automated, so the same moment gets scored the same way next time.

Hand the guide to the vendor

The point of writing the guide down is not to win the disagreement. It is to stop having the same disagreement every month. Once the behavior is named and the reasoning is written, it can be handed to whoever builds the automated score, so thenext disputed score is either earned or does not happen.

Build the guide from real disputes

The fastest way to write this guide is not to imagine every behavior in advance. It is to start from the calls where a human auditor and a score already disagree, the way the clinical lead's poor assessments already carried detailed feedback explaining why.

Each disagreement, walked through the steps above, adds one more named line to the guide. A guide built this way only contains behaviors that have already caused a real disagreement, which is exactly the set worth calibrating.

Built for an unpredictable week

A guide that only works when call volume behaves is not a guide. The same mobility company pointed out that they cannot restrict how many calls they handle, because volume in that industry tracks real incidents, not a schedule.

The same team also needs the guide, and whatever reads calls against it, to notice when a trip's status changes in the middle of a campaign rather than only at the start of one. Calibration has to hold on the week volume doubles, not only in the quiet week when someone had time to write it down.

Add it to an existing rubric

Most teams are not starting from nothing. The mobility company already scores every call against nine separate parameters. A calibration guide does not replace that structure. It defines, for each parameter, what counts as actually meeting it, the way asking about a life change and asking why it happened are two different bars sitting on the same line of a form.

Who should sit in the room

The people who need to agree are the ones who will use the disagreement afterward. That means the auditor who caught the gap, whoever owns the scoring system, and whoever will read the guide the next time a dispute comes in. A session missing one of those three produces a guide nobody downstream trusts.

Questions QA leads ask

What is a call calibration session and why does a QA team need one?

It is a working session where a human auditor and whoever owns the scoring system sit down on a specific disputed call and agree, in writing, what should have happened.

How do you know when your scoring system needs recalibrating?

The clearest signal is a specific call where the human auditor and the automated score reach opposite conclusions about the same performance. Not a general feeling that scores seem off. A score that reads tone well can still miss the clinical or factual substance entirely.

What belongs in a calibration guide besides the score itself?

A named behavior, a plain description of what a great version of that moment sounds like, and the specific call moment that first exposed the gap. Without the specific moment attached, the guide turns back into vague advice that never says more than telling someone to pay closer attention.

Should the calibration guide go to the vendor building the automated score?

Yes. The whole point of writing it down is that it can be handed to whoever builds the score, so the next call gets scored against the standard the audit team actually applies rather than a generic one.

Ready for your next calibration session?

Start from one disputed call, not a blank form. Write down what the auditor caught, what a great version of that moment sounds like, and send it to whoever scores the next one. If you want that standard applied to every call instead of a sample, see how Insight7 does it. Book a demo.