Every phone team that has ever tried call QA has a dead scorecard in a spreadsheet somewhere. It got built in an afternoon, used for three weeks, argued about twice, and quietly abandoned. The reason is almost never effort. It is that the scorecard scored opinions, and opinions do not survive being disagreed with.
A scorecard that lasts has three properties: every line is observable, every verdict carries evidence, and the whole thing is calibrated against a human before anyone is ranked by it.
One thing to settle before any of that, because it changes the answers: a scorecard is a specification, not a workflow. Who applies it — a manager sampling by hand, or a system reading every call — is a separate decision, and it moves both how long the scorecard can be and how much calibration you can afford. Where that fork matters below, it is called out.
Score behaviours, not qualities
The fastest way to kill a scorecard is a line item like “built rapport” or “was consultative”. Two managers will score the same call differently, both will be confident, and the rep will learn that the score depends on who listened.
Replace every quality with the behaviour that would have produced it. “Built rapport” becomes “used the lead’s name and referred to something they said in the first two minutes”. “Was consultative” becomes “asked at least two questions before describing the offer”. Those are checkable by anyone, including someone who was not on the call.
The test is blunt: if two people who read your checkpoint could listen to the same call and disagree about whether it happened, the checkpoint is not written yet.
Keep it to the checkpoints that move money
A twenty-line scorecard is a scorecard no human fills in. When a manager does the scoring, their attention is the binding constraint, so the useful version tracks only the handful of moments where calls actually die — which for most appointment businesses sits late in the conversation rather than early:
- The agenda set. Did the rep say what the call was for before starting it?
- Discovery before price. Did any qualifying question land before the number did?
- The value frame. Was the price delivered next to what it buys, or on its own?
- The objection answered. Was the real objection surfaced, or was the first one taken at face value?
- The locked next step. Is there a specific time on a calendar, or is there a “they’ll call back”?
Five to eight checkpoints is enough to find the leak by hand. If a ninth genuinely earns its place, add it and delete one.
Once the scoring is automated the constraint moves, and the advice inverts. Attention is no longer scarce, so the list can run longer, and conditional checkpoints — the ones that only apply to some call types — become worth carrying, because nobody is paying a cost to skip them. Teams that automate scoring usually end up with a longer rubric than the one they could work manually, not a shorter one. Treat “keep it short” as a constraint of manual QA, not a property of good scorecards.
Every verdict needs a timestamp
This is the part that decides whether reps accept the scorecard or fight it.
Feedback that says “your closing was weak this week” invites an argument, because it is a claim with nothing behind it. Feedback that says “on Tuesday’s 11:40 call, at 6:41, the lead offered to call back Thursday and the call ended there” is not an argument. It is a recording both people can open.
So the rule is: no score without a locator. Call ID and timestamp, or it does not go on the card. This costs the manager nothing at review time and removes almost all of the defensiveness, because the conversation moves from whether it happened to what to do instead.
It also disciplines whoever is marking the card. A checkpoint that cannot be pinned to a moment is usually a quality in disguise, and belongs back in the first section of this article.
Calibrate before you rank
The last step is the one most teams skip. Before the scorecard is used to compare anyone to anyone, sit two managers down with the same set of calls, have them mark them independently, then compare. Where they disagree is not noise; it is the specification you are missing.
Every call the two of them mark differently points at a checkpoint whose wording is ambiguous. Rewrite it, rescore, repeat until the disagreements are rare. Ten to twenty calls is usually enough to expose the badly worded lines — that is roughly the size of a free 10-Call Leak Snapshot, and it is a directional read, not a baseline. Publishing an agreement rate, or ranking people on one, needs a great deal more than that.
The same logic applies when a machine is doing the marking instead of a manager: an AI that has never been argued with by the person whose judgement it is meant to copy is producing confident numbers nobody has checked. Calibrate first, publish the agreement rate, then let it run. A score a manager cannot audit is a score a rep will not accept, and a rep who does not accept the score does not change the behaviour, which was the entire point.
What good looks like after a month
You should be able to answer three questions without opening a recording: which checkpoint the team fails most often, which rep is furthest from the team on that checkpoint, and which five calls a manager should sit down with this week. If the scorecard cannot answer those, it is measuring the wrong things.
And if the answers change nothing about what gets coached on Monday, the scorecard is already dead, whatever the spreadsheet says.
None of this needs us. If you would rather not build it from scratch, that is what the Calibration Sprint is: 60–90 days of your calls mined, the checkpoints your top performers already run extracted, and the scoring calibrated against your best manager’s judgement before anyone is ranked by it.