Cookies on this site
Strictly necessary cookies keep the site working. Our analytics and advertising tags — Microsoft Clarity and Google Tag Manager — stay switched off, and write no cookie, until you accept them. Privacy Policy
Cookie preferences
Choose which categories may run. Your choice is stored on this device and is remembered for six months. You can change it at any time from the “Cookie preferences” link in the footer.
Security, session integrity, your light/dark theme choice, and this cookie preference itself. The site cannot work without these, so they cannot be switched off.
Microsoft Clarity (session replay and heatmaps) and Google Analytics via Google Tag Manager. Used to see which pages help and which confuse. Off by default.
Google advertising tags via Google Tag Manager, used to measure which campaigns lead to a demo booking and to show relevant ads. Off by default.
A performance appraisal is the periodic formal evaluation of an individual's work across a defined past period. It produces two things at once: a decision about pay, promotion or continuation, and a written record of the reasoning behind that decision. It is a discrete event, distinct from the continuous management practice around it and unable to substitute for it.
An appraisal is asked to do two jobs in the same hour. One is evaluative: to reach a defensible judgment that will be used to allocate money, a title or continued employment, and to write down why. The other is developmental: to help someone see what to work on next, which requires them to talk openly about what they find hard. The second only happens when the person feels safe enough to admit a weakness, and nobody admits weakness in a conversation they know is being scored. So the evaluative purpose quietly dominates, the employee manages the impression rather than the problem, and the development portion becomes a paragraph written afterwards that neither party revisits. Employers who separate the two conversations in time usually find both improve, because each is then allowed to be what it is.
Less than its precision suggests. A rating is not a measurement of work; it is a manager's judgment, converted into a scale point, of evidence that was itself gathered by judgment. Two conversions happen before the number appears, and both are lossy. What the scale point then carries into a spreadsheet is a figure that looks like data and behaves like data downstream, feeding averages, distributions and comparisons that no one would attempt with the underlying prose. The label attached to each point does more work than the point itself, because the words on the scale tell a manager what story they are being asked to tell. A scale whose middle point reads as failure will not be used honestly, whatever the guidance says, since no manager wants to hand a person a word they will remember as an insult.
Because a surprise at appraisal is evidence that management stopped somewhere earlier. If a concern is raised for the first time in the formal conversation, the employee has been denied the period in which they could have acted on it, and they know that as soon as they hear it. What follows is an argument about the accuracy of the account rather than a discussion of the work, and the manager, who is now defending a position rather than describing one, tends to soften it. The record that results understates the problem and will be useless later. The appraisal is a summary; everything in it should already have been said while there was still time to respond. Where that discipline holds, the meeting is short and largely uncontested, which is exactly what an event of this kind should look like.
Four recur often enough to design against. Compression is the first: managers cluster their people toward the middle of the scale, because the top point invites a challenge from finance and the bottom point invites a conversation nobody wants this week. Halo is the second: one vivid strength or one memorable failure colors every other dimension, so a scale with several distinct criteria produces near-identical scores and stops carrying the separate information it was built for. The third is the leniency or severity of the individual manager, stable enough to be a property of the manager rather than of their team. The fourth is recency, and it is the most consequential: the period under review is long and memory of it is not, so the last few weeks quietly stand in for the whole.
The structural fix is cheap and unpopular: write things down when they happen. A manager who keeps a short running note arrives at the appraisal with the whole period in front of them rather than the last of it, and the note takes minutes where reconstructing it takes hours. Most [performance management software](/performance-management-software) will hold those notes, though the habit is what matters and no tool installs it. The second fix is to stop reading any single rating as a fact about a person. What is worth watching is the pattern: whether a manager's distribution sits consistently above or below their peers. Compare a manager against their own history and their peers in the same period, not against a figure someone has published as normal, because it came from a different population under different scales.
Not fairness, but comparability. A rating is only meaningful if the same words mean the same thing when two different managers write them, and there is no reason to expect that they will. Each manager has a different sample of people to compare against, a different tolerance for disappointing someone, and a different view of how hard their team's work is. Someone rated at the top of a demanding team and someone rated at the top of an easy one arrive at the same label from different places. Nothing in the form catches this: it asks each manager to judge in isolation, then treats the outputs as if they came off the same instrument. Calibration exists to expose that, by putting the judgments side by side while they can still be changed.
The session works when it asks managers to describe, not to defend. A manager who has to say out loud what a person did to earn a label, in front of peers rating comparable people, will discover quickly whether their evidence is thinner than their conviction. It works badly when it becomes a negotiation over slots, with managers trading concessions so each protects one favored person, or when a distribution is handed down in advance and the session becomes the exercise of fitting names to a shape decided elsewhere. The difference is whether the room can change a rating in either direction. Aggregating the outcomes afterwards, a natural job for [HR analytics](/hr-analytics-software), shows whether calibration is moving anything at all, or whether every manager leaves with the ratings they walked in with.
It becomes evidence. Someone will ask why one person was promoted and another was not, or why an employment ended, and the appraisal record is what answers. That audience is not the employee. It may be a successor manager who never saw the work, an internal reviewer handling a grievance, or an external party. Each reads the document without the context the writer assumed, which is why a record consisting of adjectives is worth little; what survives the journey is a description of what happened, when, and what was said about it at the time. Where a record may later support a dismissal or a contested decision, what is required of it and how long it must be kept vary by jurisdiction and change, so confirm the current position with a qualified advisor.
The failure that does the most damage is the kind record. A manager who does not want an argument writes something mild, tells themselves the person understood the tone, and moves on. Nothing in that document says there was a problem. When the same manager later needs to act, the file contradicts them: an unbroken run of satisfactory records, followed by a decision the record cannot explain. The employee reads the same file and reasonably concludes the decision came from somewhere other than their work. A record that softens a real concern is worse than no record, because no record leaves the question open while a false one answers it in the employee's favor. Storing appraisals inside the [HR software](/hr-software) makes them retrievable, but retrievability only helps if what was written was true.
Because it looks backward at a period already closed. Whatever the conversation concludes, none of it can be applied to the work it describes; that work was done, delivered and consumed while the assessment was still being formed. The employee can act on the conclusion only in the period ahead. That is not an argument against holding the event, but it bounds what the event can be asked to achieve. An organization with no other channel is asking one retrospective meeting to carry direction-setting, correction, recognition and a decision at once, and it will do the decision, because the decision has a deadline attached and the rest does not. The symptom is familiar: a calendar full of appraisals, and employees who describe feedback as something that happens to them rather than something they use.
What the event does that nothing continuous can is force a stop. Ongoing conversation is naturally local: it deals with the current problem, the thing in front of both people. It does not naturally produce a view of a whole period, it does not compare one person against another, and it produces no artifact anybody can consult later. The appraisal does all three, and those are the reasons to keep it rather than sentiment about tradition. Judged that way, the design questions become narrow and answerable. Does the period under review match a span of work anyone can actually remember and assess? Does the output feed the decisions it exists to serve, or is it collected and filed? Is the writer of the record the person who saw the work?
Pitch N Hire is an applicant tracking system built for recruiters and hiring teams. Everything on this page β sourcing, screening, interviewing, offers β runs in one pipeline.
Free for 1 user Β· No credit card Β· Talk to a real hiring expert
Pitch N Hire unifies sourcing, screening and hiring decisions on one AI-native platform. Book a quick demo on your real roles.
Prefer to talk? Book a demo Β· Talk to sales Β· View pricing
Free 1-user plan Β· No credit card Β· Talk to a real hiring expert
See your true cost-per-hire and how much Pitch N Hire could save you β our free Recruitment ROI Calculator gives you the numbers in under a minute. No signup required.
Open the free ROI calculatorPrefer a tailored walkthrough on your real roles? Drop your work email:
β Free 1-user plan Β· No spam Β· Talk to a real hiring expert