Give every rating a written anchor describing what evidence earns it, then have interviewers score the same real candidate independently and compare in a group. Differences almost always come from different standards rather than different observations. Review scoring patterns per interviewer each quarter and recalibrate anyone whose ratings drift consistently high or low.
Usually because they are answering different questions. One interviewer scores against the role as defined, another against themselves at that stage of their career, a third against the last person they interviewed. Without written anchors, a rating of four means whatever the individual thinks it means. Two other causes appear regularly. Interviewers assess different competencies in the same conversation and then compare overall impressions, which are not comparable. And some interviewers rate the interaction rather than the evidence, so an articulate candidate with thin examples outscores a quieter one with strong ones. All three are structural problems, and none is solved by telling people to be more objective. Naming which of the three is happening on your team is the first useful step, because the remedies are different.
A short description of the evidence that earns each score, written for that specific competency and role level. For a systems design competency, the top anchor might describe someone who designed a comparable system, explained the trade-offs they rejected, and could say what they would change now. The bottom anchor describes someone who described a design they did not make decisions inside. Write the middle points too, since that is where most candidates land and where disagreement concentrates. Keep it to a line per level. Anchors written for the specific role beat generic five-point scales, and they are what makes a scorecard something more than a form. Store them with the [structured interview questions](/interview-questions) each interviewer uses.
Take one real recorded or recent candidate, have every interviewer score independently before any discussion, then reveal all scores at once. Start with the widest gap and ask each person for the specific evidence behind their rating. The conversation is the training, and it works because the disagreement is concrete rather than hypothetical. Expect the first session to reveal that people were not assessing the same thing at all. Run these quarterly, and always before a hiring surge or when new interviewers join. Independent scoring before discussion matters more than anything else in the format, because once a senior voice speaks first, the rest of the scores quietly converge on it.
Look at each interviewer's rating distribution across a quarter. Someone who has never given a low score is not seeing better candidates than everyone else, and someone who rejects nearly everyone is applying a private standard. Compare each interviewer's ratings with the eventual outcome as well: whose positive ratings preceded hires that worked out, and whose reservations were repeatedly overruled and repeatedly correct. That second pattern is worth acting on. Most [recruitment analytics](/recruitment-analytics-software) tools can produce these distributions from scorecard data you already collect. Share the numbers privately, framed as calibration rather than performance, otherwise interviewers start scoring to the middle to avoid attention. Recalibrate anyone whose distribution sits well away from the group, and do it with examples rather than with the chart alone.
Get a personalized walkthrough of Pitch N Hire on your own roles and workflow. No slides, no obligation.
Prefer to talk? Book a demo · View pricing
Free 1-user plan · No credit card · Talk to a real hiring expert
See your true cost-per-hire and how much Pitch N Hire could save you — our free Recruitment ROI Calculator gives you the numbers in under a minute. No signup required.
Open the free ROI calculatorPrefer a tailored walkthrough on your real roles? Drop your work email:
★ Free 1-user plan · No spam · Talk to a real hiring expert