Interviewing & Assessment

Interview Calibration

Interview calibration is the practice of getting interviewers to apply one standard to the same evidence, so a candidate's rating does not depend on who happened to be assigned. It runs on shared rating definitions, joint scoring of the same recorded interview, shadow sessions, and periodic review of how each interviewer's scores are distributed.

What does interview calibration look like in practice?

Four mechanisms carry most of the work. Shadowing puts a new interviewer in the room to watch an experienced one, and reverse shadowing flips it so the experienced interviewer observes and critiques. Joint scoring takes one recorded interview, has six people rate it independently, then compares the ratings and argues about the gaps until the group agrees what the evidence supported. Rating definitions turn each point on the scale into observable behaviour rather than a feeling. Distribution review looks across every interviewer's scores over a quarter to find the person who never gives below a four and the person who never gives above a three. Run the first two whenever someone joins the panel, the second two on a recurring cadence. Recorded sessions and structured scorecards held in an [applicant tracking system](/ats) make the last two possible at all.

How is calibration different from interviewer training?

Training teaches the mechanics: how to ask a behavioural question, how to probe for a specific example, what to write down, which subjects to avoid. Calibration assumes the mechanics are already there and works on where the bar sits. You can have a fully trained panel whose members land a full point apart on the same candidate, because one of them privately expects senior-level depth from a mid-level hire and has never said so out loud. Training fixes a skills problem. Calibration fixes an alignment problem, and it never finishes, because standards drift as teams grow, as the market moves and as new interviewers join. Treat them as separate programmes with separate owners. Teams that train once and consider the job done keep the variance they set out to remove.

How do you define what a 3 versus a 4 actually looks like?

Write the scale in observable behaviour, then attach real examples. A four on cross-team communication might read: explained a technical trade-off to a non-technical stakeholder, named the constraint being optimised for, and described how the other team responded. A three might read: described the outcome clearly but could not explain the trade-off when pushed. Vague anchors such as strong or meets expectations produce variance, because every interviewer fills them in from personal history. Once anchors exist, test them. Give the panel three anonymised answer transcripts, ask everyone to rate, and rewrite any anchor where the group splits. Keep the scale short. Four points with an explicit hire and no-hire boundary beats a ten-point scale nobody can distinguish between.

What do score distributions tell you about your interviewers?

They surface the two rater problems that no amount of discussion catches. The lenient interviewer, whose median sits high, passes candidates the rest of the panel would reject, and that shows up later as regretted hires from loops they sat on. The severe interviewer rejects people other panels would have hired, and that cost stays invisible because the candidate simply goes elsewhere. Pull each interviewer's ratings over a meaningful number of interviews, compare medians and spread, and talk to the outliers with the data in front of you. Some variance is legitimate, since interviewers do not see identical candidate pools. Persistent variance is not. [Recruitment analytics](/recruitment-analytics-software) built on scorecard data turns this into a routine report instead of a research project.

See how Pitch N Hire handles interview calibration on your roles

FAQ

Interview Calibration — FAQs

How often should a hiring team recalibrate? +
Quarterly for panels that interview regularly, and immediately after any change that moves the bar: a new level definition, a new compensation band, a shift from senior to junior hiring. Recalibrate too when a panel's hires start diverging from performance reviews. Interviewers who go months without a session drift, so run a short refresher before returning them to active loops.
What is reverse shadowing? +
Reverse shadowing is the second half of the apprenticeship. In a shadow session the trainee watches an experienced interviewer run the interview. In reverse shadowing the trainee runs it while the experienced interviewer observes silently, then compares scores and gives feedback afterwards. Most panels require at least one of each, plus one matching score, before a new interviewer runs a session alone.
Does calibration mean everyone has to agree on a candidate? +
No. Disagreement about what a candidate demonstrated is useful and often the most valuable part of a debrief. What calibration removes is disagreement about the standard itself, where two people watched the same answer and one called it strong only because their internal bar sits lower. Aligned bars, independent judgements. That combination is the goal.
Can you calibrate without recording interviews? +
Yes, though it takes longer. Use written transcripts of past answers, or run a live mock interview with a volunteer while the panel scores from the room. You can also calibrate retrospectively by pulling scorecards on candidates whose outcomes you now know. A shared bank of [structured interview questions](/interview-questions) helps, since calibration is easier when everyone asked the same thing.
Built for recruiters & hiring teams

See Interview Calibration in action

Pitch N Hire unifies sourcing, screening and hiring decisions on one AI-native platform. Book a quick demo on your real roles.

Prefer to talk? Book a demo · View pricing

Free 1-user plan · No credit card · Talk to a real hiring expert

One Hiring Infrastructure.
Zero Tool Chaos.

Demos are consultative. We respect privacy and enterprise
governance. No lock-ins.

Start free Book demo