Back to the Journal

Essay · Sales Coaching

How to build a sales call scorecard

(And keep using it past month three.)

Anna SivénFounder & CEO, Velisi

Most sales call scorecards are built in one enthusiastic afternoon and quietly abandoned before the quarter ends. The building is not the hard part. Sitting down every week to score calls against a rubric nobody has calibrated, while the pipeline is asking for attention, is the hard part.

So the useful question is not what a complete sales call scorecard would contain. It is what the smallest one is that would still tell you something you did not already know, and whether you could defend a score when a rep disagrees with it. Both of those are design decisions, and both are usually skipped.

What a sales call scorecard is for

A scorecard converts a judgement about a call into a number that can be compared: this rep against their own baseline, this month against last, this call against the standard you claim to hold. That is all it does, and it is worth saying plainly, because scorecards get asked to do two other jobs they are bad at.

They are not a performance-management instrument. The moment a score decides compensation, the scores stop describing calls and start describing the relationship between rep and reviewer. And they are not a substitute for listening: a 3 tells a rep nothing they can act on without the sentence explaining why it was a 3.

The version that works is narrow. A shared definition of what good looks like, applied consistently enough that movement over time means something. The coaching still happens in conversation. The scorecard only decides what the conversation is about.

How to build a sales call scorecard

Start with four criteria, not fourteen

The instinct is to cover the whole call: rapport, discovery, qualification, framing, objection handling, next steps, tone, pace. Fourteen criteria produce a form long enough that filling it in properly competes with everything else in a manager's afternoon, which means it stops being filled in properly rather quickly.

Pick four, and pick them from where your deals actually fail. If deals stall after the demo, the criteria are about discovery depth and qualification, not rapport. If deals go quiet after price, the criteria are about how price was framed and what happened in the ten seconds after the objection landed. A scorecard that covers everything prioritises nothing.

Write anchors, not adjectives

"Strong discovery" is not a criterion. It is an opinion with a number attached. An anchor describes what a scorer would have to observe.

For a discovery criterion: 1 means no question about business impact was asked at all; 3 means impact was raised once and never quantified; 5 means impact was quantified and the rep repeated it back before moving on. Two reviewers reading that will land within a point of each other. Two reviewers reading "strong discovery" will not, and neither will the same reviewer on a Monday and a Friday.

Anchors are most of the real work in building a scorecard, and the part most often skipped. They are also what makes a score defensible when a rep pushes back, which they will and should.

Decide the sample before you start

Four or five calls per rep per month, drawn at random rather than chosen. Choosing reintroduces the bias you are trying to remove: managers pick memorable calls, and memorable usually means bad.

Random sampling also protects the exercise from the obvious objection, which is that five calls cannot represent a rep's month. They cannot. What they can support is a comparison against the same rep's sample last month, and that is the only comparison a small sample honestly carries.

Calibrate before you trust a number

Have two reviewers score the same three calls independently, then compare. The first round is usually worse than anyone expects, and the widest gaps tend to sit on the criteria everyone believed were obvious. That gap is information about the anchors rather than about the reviewers, and rewriting the anchor is the fix.

Until you have run one calibration round, treat the scores as private working notes. A number shared with a rep before it is stable costs more trust than it buys, and trust is the input the whole exercise runs on.

What a scorecard cannot tell you

It cannot tell you whether the call was winnable. Some buyers were never buying, and a well-run call with a bad-fit prospect scores high and closes nothing. Scorecards measure execution against a standard, not outcome, and confusing the two produces the familiar, demoralising conversation in which a rep is marked down for a deal that was lost upstream of them.

It also cannot tell you the standard is right. If the rubric encodes a playbook that does not work in your market, consistent scoring will produce consistent mediocrity, measured precisely. The check is to compare high-scoring calls against won deals every so often and ask whether they are the same calls. Often enough they are not, and noticing that is the most valuable thing a scorecard ever produces.

And it will drift. Reviewers get more generous across a year without deciding to. Recalibrating twice a year is the running cost of keeping a scorecard comparable to itself, which is the same measurement discipline described in how to measure sales coaching effectiveness.

Making it survive month three

Every scorecard dies the same way. The effort per call is fine at the start and unaffordable a couple of months in, because scoring competes with the pipeline every single time and the pipeline always wins the argument.

Three things extend its life. Keep it to four criteria and one line of comment each. Score on a fixed day rather than when there is time, because there is never time. And connect the scores to the one behavior you are currently coaching, so the exercise has a visible use instead of being a filing task.

The underlying failure is the one described in why post-call coaching fails: the ritual outlives the effect. A scorecard nobody reads is a slower version of a call recording nobody opens, and it costs more to produce. The same applies to the raw material — pattern across calls is what you need, not volume, which is the distinction drawn in what sales call analytics is.

Where Velisi sits in this

Velisi listens to live calls and surfaces the nudge during the conversation rather than after it, lets reps practise against realistic AI buyers before the real call, and tracks skill gaps and improvement per rep so a manager sees the pattern across a team instead of a folder of recordings. That lowers the cost of the two expensive parts of everything above: noticing what happened on a call, and seeing whether it changed. Velisi is currently accepting applications from selected sales teams, and the pieces are laid out on the AI sales coaching platform page.

It does not remove the need for judgement. Somebody still has to decide what good looks like on your calls, write the anchors, and defend them in a conversation with a rep who disagrees. That part stays yours, and it should.

— Anna

Frequently asked questions

What is a sales call scorecard?

A short rubric used to score a sales call against a defined standard, usually three to five criteria with described levels. Its purpose is comparison over time — a rep against their own baseline — rather than a verdict on any single call.

How many criteria should a sales call scorecard have?

Four is a sensible ceiling. Beyond that the time per call rises past what a manager sustains, and the exercise is usually abandoned within a quarter. Choose the criteria from the stage where your deals actually fail, not from a general model of a good call.

Who should score sales calls?

Whoever coaches the rep, with a second reviewer used for calibration. Reps scoring their own calls is useful for self-awareness and close to useless for comparison, because the standard drifts differently for each person.

Should scorecard results affect compensation?

No. Once a score carries money it stops describing the call and starts describing a negotiation between rep and reviewer. Keep scorecards inside coaching, where a low score is information rather than a penalty.

How many calls per rep should you score each month?

Four or five, sampled at random rather than chosen. It is not enough to characterise a rep's month, and it is enough to compare against that same rep's sample from the previous month, which is the comparison a small sample honestly supports.

How often should a scorecard be recalibrated?

Twice a year, plus any time two reviewers land more than a point apart on the same call. Scoring drifts gently towards generosity, and drift makes month-to-month comparison meaningless without anyone noticing it happen.

AS

Anna Sivén

Founder & CEO, Velisi

Applications open

Score the call, then coach the person.

Velisi is currently accepting applications from selected sales teams.

Apply for consideration →