Interviewing

How do I calibrate interviewers who rate candidates differently?

Give every rating a written anchor describing what evidence earns it, then have interviewers score the same real candidate independently and compare in a group. Differences almost always come from different standards rather than different observations. Review scoring patterns per interviewer each quarter and recalibrate anyone whose ratings drift consistently high or low.

Why do interviewers score the same candidate differently?

Usually because they are answering different questions. One interviewer scores against the role as defined, another against themselves at that stage of their career, a third against the last person they interviewed. Without written anchors, a rating of four means whatever the individual thinks it means. Two other causes appear regularly. Interviewers assess different competencies in the same conversation and then compare overall impressions, which are not comparable. And some interviewers rate the interaction rather than the evidence, so an articulate candidate with thin examples outscores a quieter one with strong ones. All three are structural problems, and none is solved by telling people to be more objective. Naming which of the three is happening on your team is the first useful step, because the remedies are different.

What does a good rating anchor look like?

A short description of the evidence that earns each score, written for that specific competency and role level. For a systems design competency, the top anchor might describe someone who designed a comparable system, explained the trade-offs they rejected, and could say what they would change now. The bottom anchor describes someone who described a design they did not make decisions inside. Write the middle points too, since that is where most candidates land and where disagreement concentrates. Keep it to a line per level. Anchors written for the specific role beat generic five-point scales, and they are what makes a scorecard something more than a form. Store them with the [structured interview questions](/interview-questions) each interviewer uses.

How do you run a calibration session?

Take one real recorded or recent candidate, have every interviewer score independently before any discussion, then reveal all scores at once. Start with the widest gap and ask each person for the specific evidence behind their rating. The conversation is the training, and it works because the disagreement is concrete rather than hypothetical. Expect the first session to reveal that people were not assessing the same thing at all. Run these quarterly, and always before a hiring surge or when new interviewers join. Independent scoring before discussion matters more than anything else in the format, because once a senior voice speaks first, the rest of the scores quietly converge on it.

How do you spot drift over time?

Look at each interviewer's rating distribution across a quarter. Someone who has never given a low score is not seeing better candidates than everyone else, and someone who rejects nearly everyone is applying a private standard. Compare each interviewer's ratings with the eventual outcome as well: whose positive ratings preceded hires that worked out, and whose reservations were repeatedly overruled and repeatedly correct. That second pattern is worth acting on. Most [recruitment analytics](/recruitment-analytics-software) tools can produce these distributions from scorecard data you already collect. Share the numbers privately, framed as calibration rather than performance, otherwise interviewers start scoring to the middle to avoid attention. Recalibrate anyone whose distribution sits well away from the group, and do it with examples rather than with the chart alone.

Want Pitch N Hire to handle this for your team?

Related glossary terms

Next step

FAQ

Frequently asked questions

What do you do when two interviewers strongly disagree? +
Ask both for the specific evidence rather than the conclusion, and check whether they assessed the same competency. Genuine disagreement on the same evidence usually means the bar for that competency is undefined, and that is the thing to fix. If it stays unresolved, one more targeted interview on that single competency answers it faster than another debate.
Should interviewers see each other's feedback before the debrief? +
No. Submitted independently and revealed together is the rule, because prior feedback anchors judgement, particularly when it comes from someone senior. Most applicant tracking systems can lock feedback visibility until submission. Turn that setting on, since asking people to avoid reading is far less reliable than removing the option.
How many interviewers should assess the same competency? +
Two is a reasonable balance for competencies that carry real risk, one is enough for the rest. Assigning four people to the same competency produces overlapping evidence and a longer debrief without improving the decision. Coverage across all the competencies matters more than redundancy on your favourite one.
Does calibration slow hiring down? +
The sessions cost an hour each quarter and typically shorten debriefs considerably, because most debrief time is spent reconciling scores that meant different things. Teams with anchored scorecards reach decisions faster and with fewer repeat interviews, which is where the time comes back. The first session is the expensive one, and later ones run shorter because the standard is already written down.
Built for recruiters & hiring teams

See how much faster your team could hire

Get a personalized walkthrough of Pitch N Hire on your own roles and workflow. No slides, no obligation.

Prefer to talk? Book a demo · View pricing

Free 1-user plan · No credit card · Talk to a real hiring expert

One Hiring Infrastructure.
Zero Tool Chaos.

Demos are consultative. We respect privacy and enterprise
governance. No lock-ins.

Start free Book demo