Cookies on this site
Strictly necessary cookies keep the site working. Our analytics and advertising tags — Microsoft Clarity and Google Tag Manager — stay switched off, and write no cookie, until you accept them. Privacy Policy
Cookie preferences
Choose which categories may run. Your choice is stored on this device and is remembered for six months. You can change it at any time from the “Cookie preferences” link in the footer.
Security, session integrity, your light/dark theme choice, and this cookie preference itself. The site cannot work without these, so they cannot be switched off.
Microsoft Clarity (session replay and heatmaps) and Google Analytics via Google Tag Manager. Used to see which pages help and which confuse. Off by default.
Google advertising tags via Google Tag Manager, used to measure which campaigns lead to a demo booking and to show relevant ads. Off by default.
Give every rating a written anchor describing what evidence earns it, then have interviewers score the same real candidate independently and compare in a group. Differences almost always come from different standards rather than different observations. Review scoring patterns per interviewer each quarter and recalibrate anyone whose ratings drift consistently high or low.
Usually because they are answering different questions. One interviewer scores against the role as defined, another against themselves at that stage of their career, a third against the last person they interviewed. Without written anchors, a rating of four means whatever the individual thinks it means. Two other causes appear regularly. Interviewers assess different competencies in the same conversation and then compare overall impressions, which are not comparable. And some interviewers rate the interaction rather than the evidence, so an articulate candidate with thin examples outscores a quieter one with strong ones. All three are structural problems, and none is solved by telling people to be more objective. Naming which of the three is happening on your team is the first useful step, because the remedies are different.
A short description of the evidence that earns each score, written for that specific competency and role level. For a systems design competency, the top anchor might describe someone who designed a comparable system, explained the trade-offs they rejected, and could say what they would change now. The bottom anchor describes someone who described a design they did not make decisions inside. Write the middle points too, since that is where most candidates land and where disagreement concentrates. Keep it to a line per level. Anchors written for the specific role beat generic five-point scales, and they are what makes a scorecard something more than a form. Store them with the structured interview questions each interviewer uses.
Take one real recorded or recent candidate, have every interviewer score independently before any discussion, then reveal all scores at once. Start with the widest gap and ask each person for the specific evidence behind their rating. The conversation is the training, and it works because the disagreement is concrete rather than hypothetical. Expect the first session to reveal that people were not assessing the same thing at all. Run these quarterly, and always before a hiring surge or when new interviewers join. Independent scoring before discussion matters more than anything else in the format, because once a senior voice speaks first, the rest of the scores quietly converge on it.
Look at each interviewer's rating distribution across a quarter. Someone who has never given a low score is not seeing better candidates than everyone else, and someone who rejects nearly everyone is applying a private standard. Compare each interviewer's ratings with the eventual outcome as well: whose positive ratings preceded hires that worked out, and whose reservations were repeatedly overruled and repeatedly correct. That second pattern is worth acting on. Most recruitment analytics tools can produce these distributions from scorecard data you already collect. Share the numbers privately, framed as calibration rather than performance, otherwise interviewers start scoring to the middle to avoid attention. Recalibrate anyone whose distribution sits well away from the group, and do it with examples rather than with the chart alone.
Pitch N Hire is an applicant tracking system built for recruiters and hiring teams. If this answer described something you want to run properly, the ATS is where it lives.
Free for 1 user · No credit card · Talk to a real hiring expert
Get a personalized walkthrough of Pitch N Hire on your own roles and workflow. No slides, no obligation.
Prefer to talk? Book a demo · Talk to sales · View pricing
Free 1-user plan · No credit card · Talk to a real hiring expert
See your true cost-per-hire and how much Pitch N Hire could save you — our free Recruitment ROI Calculator gives you the numbers in under a minute. No signup required.
Open the free ROI calculatorPrefer a tailored walkthrough on your real roles? Drop your work email:
★ Free 1-user plan · No spam · Talk to a real hiring expert
Free 1-user plan · No credit card · Real person replies in 1 business day