How to Score a Sales Roleplay (And Test Your Rubric)
Score a sales roleplay on 3–5 anchored behaviours, then run the calibration check that tells you whether two managers grading the same run would agree.
In short
Score three to five observable behaviours on a short anchored scale, from the recording. Then calibrate: two graders score the same eight runs, you compute agreement, and you fix the rubric before you use the numbers for anything that matters.
Score a sales roleplay on three to five observable behaviours with written anchors, graded from the recording. Then test the rubric: two managers score the same eight runs, and you check they agree before trusting any number it produces.
Every guide to this topic stops at the first sentence. They tell you to pick criteria, align them to your methodology, and use a 1–5 scale. All true, all insufficient. A rubric is a measuring instrument, and nobody publishes the step where you check the instrument is measuring anything.
That step is below, as a protocol you can run in about 90 minutes.
What should a sales roleplay rubric actually score?
Behaviours a grader can point at in the recording, not qualities they have to judge. "Asked a question before making a product claim" is scoreable. "Built rapport" is not — it is a conclusion, and two graders will reach different ones from the same tape.
The test for a criterion is whether you could timestamp it. If a grader can write "04:12" next to the score, the criterion is observable. If the only defence of a score is "that's how it came across," the criterion is an opinion wearing a number.
Keep the list short. Three to five criteria per drill, each scored 0–2. A long rubric feels rigorous and grades worse, because a grader working through twenty line items stops watching the recording and starts filling in a form. The public sales-competition rubrics show what the long version looks like: the Southern Utah University competition rubric scores 24 line items across a 290-point scale, which works for a one-off contest with trained judges and does not survive a manager grading four runs on a Thursday.
| Instead of | Score this | Because |
|---|---|---|
| Built rapport | Referenced something specific to this buyer's situation in the first two minutes | One is a verdict, the other is a moment you can find |
| Handled the objection well | Asked a clarifying question before responding to the objection | "Well" is where grader disagreement lives |
| Good discovery | Asked at least three open questions before presenting | Countable, and it survives being argued about |
| Strong close | Secured a specific next step with a date | Either the date is in the recording or it is not |
| Confident delivery | Did not concede on price before the buyer named a number | Confidence is inferred; the concession is observable |
Why do most sales roleplay rubrics produce meaningless scores?
Three failure modes, all of them documented in the performance-rating literature for a century: halo, leniency, and unanchored scale points. Each one is detectable from your own scoring data, which is why the calibration protocol below is worth the 90 minutes.
Halo is the oldest of the three. Thorndike named it in 1920 after finding that raters' judgements of separate qualities correlated far more tightly than the qualities themselves plausibly did — a rater who likes the person marks everything up. In a roleplay debrief it looks like a rep scoring 2 on all four criteria, run after run. If your criteria never disagree with each other, you are not scoring four things.
Leniency is the ceiling problem. When most scores land in the top band the rubric has stopped discriminating, and a rep who is genuinely stuck looks the same as a rep who is ready. This is the failure that makes readiness dashboards worthless, and it is invisible unless you look at the distribution rather than the individual scores.
Unanchored scale points are the easiest to fix and the most common. If your rubric says 1–5 and nothing else, every grader invents their own 3. Smith and Kendall solved this in 1963 with behaviourally anchored rating scales: write a concrete example of behaviour at each scale point, and check that independent people can sort the examples back to the right point. That check is the part everyone skips.
How do you write anchors two graders will agree on?
Collect real moments from recordings, write one anchor per scale point in the words a grader would use, then have a second person sort the anchors back into their levels blind. Anything they misfile is written badly and gets rewritten.
- Pull six to ten clips of the behaviour from real roleplay recordings — a mix of good, bad, and ambiguous.
- Sort them into three piles: clearly meets the bar, partially meets it, does not.
- Write one sentence per pile describing what the rep did, using verbs. Those three sentences are your 2, 1 and 0 anchors.
- Hand the anchors and the clips to someone who was not in the room, without the pile labels, and ask them to match each clip to an anchor.
- Rewrite any anchor that pulled the wrong clip. Repeat once. Two passes is usually enough.
This is a lighter version of the retranslation procedure from the original 1963 paper, and it is the difference between a rubric that reads well and a rubric that scores consistently. Skipping it is what produces a document everybody praises in the meeting and nobody applies the same way afterwards.
How do you know your rubric actually works?
Run four checks on eight double-scored recordings: agreement between graders, score spread, criterion independence, and drift over time. A rubric that fails any of them is producing numbers you should not report upward.
This is the protocol. Two graders, eight recordings from your existing roleplay library, scored independently — no talking, no shared screen. It takes about 90 minutes including the debrief, and you should run it before the rubric drives anything consequential and again once a quarter.
| Check | How to compute it | Threshold I would use | If it fails |
|---|---|---|---|
| 1. Grader agreement | Percent of criterion-level scores where both graders gave the same number, then Cohen's kappa: k = (po − pe) ÷ (1 − pe), where po is observed agreement and pe is the agreement you would expect by chance | Kappa above 0.6 | Rewrite the anchors on the criteria that disagreed most, then re-run |
| 2. Score spread | Share of scores sitting in the top band across all criteria and all runs | Under 60% in the top band | The bar is too low. Raise the anchor for the top band rather than telling graders to be harsher |
| 3. Criterion independence | How often a rep received identical scores on every criterion in the same run | Under a third of runs | Halo. Score one criterion across all eight recordings before moving to the next, rather than scoring a run all at once |
| 4. Drift | Re-score two recordings you scored last quarter and compare to the original scores | No systematic movement in one direction | Standards have slid. Re-run the anchor sort, do not just remind people |
On kappa: the 0.6 line comes from the interpretation bands Landis and Koch published in 1977, where 0.61–0.80 is described as substantial agreement. Those bands are a convention that stuck, not a statistical truth, and they were written for medical categorical data rather than sales debriefs. Use them as a starting line and tighten if your decisions are consequential. If your scale is ordinal, linear-weighted kappa is the more honest version, because it treats a 2-versus-1 disagreement as smaller than a 2-versus-0.
The uncomfortable part of this protocol is that most rubrics fail check 1 the first time. That is the expected result, not a sign your managers are careless. Rater training research has been consistent on this point for decades: frame-of-reference training, where raters practise on shared examples and discuss the discrepancies, is the intervention with the strongest track record for improving rating accuracy. The double-scoring session is that training, run as a side effect of the audit.
What do you do when two graders disagree?
Do not average them. Find the timestamp each grader was scoring, and decide whether the disagreement is about the evidence or the anchor. Evidence disagreements are a coaching issue; anchor disagreements are a rubric defect, and only one of those is fixed by talking about it.
- Both graders point to the same moment and score it differently: the anchor is ambiguous. Rewrite it.
- The graders point to different moments: one of them missed something. Add a note to the rubric about where in the call the behaviour is expected to appear.
- One grader scored from an impression with no timestamp: that score does not count. Re-score from the recording.
- Both are right and the behaviour genuinely sat between two anchors: your scale has too few points for this criterion, or the criterion is really two criteria.
Averaging feels fair and destroys the signal. The whole point of double-scoring is to surface where the instrument is broken, and a mean quietly hides exactly the cases you needed to see.
Does any of this change when an AI scores the roleplay?
The checks matter more, not less. An automated scorer is perfectly consistent with itself, which means the usual reliability check — do two graders agree — silently passes and tells you nothing. What you have to test instead is whether the machine agrees with your managers.
Run the same protocol with the AI as one of the two graders and an experienced manager as the other. You are looking for the same four things, plus one: the criteria where the machine and the human systematically part ways. Those are usually the criteria that depend on context the scorer does not have — whether the objection was real, whether that concession was defensible given the account.
Treat that as a scoping decision rather than a defect. Automate the criteria that hold up under the check, keep the ones that do not for the human debrief, and be explicit with your team about which is which. A team that knows the machine grades question count and the manager grades judgement will trust both. A team told the AI grades everything will stop trusting either the first time a score looks wrong.
If you want to run the protocol against your own recordings, you need scenarios that hold a consistent difficulty first — otherwise you are measuring the buyer, not the rep. The buyer briefs in our sales objection roleplay scenarios are built for that, and the trade-offs between AI and human practice are covered in AI roleplay training. Voxento runs roleplays by voice: you can try a live conversation without signing up, and current plans are on the pricing page.
How often should you recalibrate?
Quarterly, and after any change to the rubric, the scenario library, or who does the grading. Calibration decays with all three, and it decays silently — nothing in your dashboard changes appearance when the numbers stop meaning what they used to.
The quarterly run is cheaper than the first one. You already have anchors, you already have a scoring sheet, and check 4 only needs the two recordings you kept from last time. Keep those two recordings deliberately. They are your reference standard, and a scoring system without one has no way to notice it has moved.
Frequently asked questions
- What scale should I use to score a sales roleplay?
- 0–2 per criterion, with a written anchor at each point. Three points force a decision and are easy to anchor. A 1–5 scale needs five defensible anchors per criterion, and in practice graders use three of them anyway, which just adds noise.
- How many recordings do I need to calibrate a rubric?
- Eight double-scored runs is enough to expose a badly written anchor, which is what the first pass is for. If you are using scores for compensation or promotion decisions, that sample is too small and you should be talking to someone who does measurement for a living.
- Should managers score live or from the recording?
- From the recording. Live scoring means the grader is also running the session, and it makes check 1 impossible because nobody else saw the same run. Score from the tape, timestamp each score, and the disagreements become findable.
- Can a roleplay score predict real sales performance?
- Not on its own, and be sceptical of anyone claiming otherwise without showing the correlation on their own data. What a calibrated rubric gives you is a consistent measure of whether a specific behaviour is present. Whether that behaviour drives your win rate is a separate question you have to answer with your own pipeline data.
- What is an acceptable level of agreement between graders?
- Kappa above 0.6 is a reasonable working line, following the Landis and Koch interpretation bands. Below 0.4 the rubric is not usable and the anchors need rewriting. Above 0.8 is unusual on judgement-heavy criteria and is worth a second look for leniency.
Sources
- Thorndike, E. L. (1920). A constant error in psychological ratings. Journal of Applied Psychology, 4(1), 25–29.
- Smith, P. C., & Kendall, L. M. (1963). Retranslation of expectations: An approach to the construction of unambiguous anchors for rating scales. Journal of Applied Psychology, 47(2), 149–155.
- Landis, J. R., & Koch, G. G. (1977). The measurement of observer agreement for categorical data. Biometrics, 33(1), 159–174.
- Woehr, D. J., & Huffcutt, A. I. (1994). Rater training for performance appraisal: A quantitative review. Journal of Occupational and Organizational Psychology, 67(3), 189–205.
- Roch, S. G., Woehr, D. J., Mishra, V., & Kieszczynska, U. (2012). Rater training revisited: An updated meta-analytic review of frame-of-reference training. Journal of Occupational and Organizational Psychology, 85(2), 370–395.
- Southern Utah University — Thunder Sales Competition Role Play Rubric (public PDF, reviewed August 2026)
- PitchMonster — Sales Role-Play Training: Best Practices + Checklist, reviewed August 2026
- Mindtickle — AI Sales Role Play: 2026 Guide for Enablement Leaders, reviewed August 2026
Written by
Muhammad Amin — Co-founder, Voxento
I co-founded Voxento and build the platform. I work directly with the schools and training teams running observations and AI roleplay on it, which is where most of what I write here comes from.
Related reading
- How to Run a Paid Training Community Profitably
Community pricing guides stop at what members will pay. Here is the cost side: platform fees, metered capacity, and contribution margin per tier.
- Voice AI vs Text Chatbots for Training
Voice practice is metered per minute and capped by concurrent sessions. Text is not. What that changes about how you design and budget a program.
- Sales Objection Roleplay Scenarios That Push Back
Most roleplay guides script the rep's lines. These 10 scenarios brief the buyer instead: hidden motive, concession rules, fail triggers, and a scoring rubric.
See AI roleplay training in action.
Build courses, run AI voice roleplays, score performance automatically, and run a branded community — in one platform.