Calibrating Ratings Fairly Across a Team

Rating drift rarely looks like bias from the inside. It looks like being a little generous with the person who reminds you of yourself, a little harsh on the one whose wins you saw less of, and calling both of those "just my read on the year." Here are the checks worth running against your own ratings before they go anywhere near a calibration room.

Key Takeaways

  • Rate against what is written down, not what you remember. Memory is reliably kinder to the people you interact with most.
  • A forced distribution you resent fighting is doing its job. One that feels effortless usually means you skipped the fight.
  • The last six weeks of the year should not carry more weight in a rating than the other forty-six.
  • Write the reasoning down before you defend the number out loud. The writing is where sloppy logic actually shows itself.

The principles

These aren't steps to complete in order - they're the things worth checking yourself against, before and during the work. Read them once, then keep this page open while you go.

  1. Rate against evidence, not memory

    Memory is not neutral. It over-weights whoever you spoke to most, whichever win was most recent, and whoever's working style most resembles your own. None of that is a fair basis for a number that affects someone's pay.

    • Pull the actual logged achievements, 1:1 notes, or peer feedback before you write a rating, not after you've already landed on one.
    • For anyone you rate from a general impression rather than specific evidence, treat that as a flag to go find the evidence, not a shortcut to trust the impression.
    • Notice who you can list five concrete wins for without checking anything, and who you can't. That gap is usually about your own visibility into their work, not their output.
    • If you keep an ongoing record like Perform Review's Achievement Log for your reports, this is exactly the evidence to pull from instead of relying on memory.
  2. Force the distribution deliberately

    Whether or not your company has a formal curve, most teams naturally cluster ratings in the middle unless someone actively pushes back on that. Deliberately checking your own spread catches the drift a first pass misses.

    • Look at your full list sorted by rating before you finalize anything. A distribution that is suspiciously flat is worth a second look, not a sign everyone genuinely performed the same.
    • Ask yourself specifically who is at the top and why, and who is at the bottom and why. "I couldn't decide so I put most people in the middle" is not a rating decision, it's the absence of one.
    • If your top and bottom ratings are the same people every single cycle with no real movement between them, that's worth questioning too. It can mean real static performance, or it can mean you've stopped actually looking.
  3. Watch for recency bias

    The last six to eight weeks before a review cycle are disproportionately vivid in anyone's memory, evidence-based process or not. A rating built mostly from that window isn't rating the year - it's rating whichever month happened to end it.

    • Check whether the evidence behind a rating is actually spread across the period, not clustered in the final stretch.
    • For a big recent win or a recent stumble, ask deliberately: does this actually change the pattern across the year, or is it just the freshest thing in your head?
    • If someone had a strong first half and a quiet second half, or the reverse, make sure both halves are genuinely represented in the number, not just whichever one you remember more clearly.
  4. Cross-check for pattern bias

    Bias in ratings is rarely a single decision you'd catch in the moment. It's a pattern that only becomes visible once you look at the whole team side by side rather than one review at a time.

    • Once ratings are drafted, look at them grouped by tenure, team, working style, or anything else that might correlate. Not to accuse yourself of anything, just to see what the numbers actually show.
    • If a pattern shows up, ask whether it reflects a real, explainable difference in output, or whether it's easier to explain as something about how each person communicates their work to you.
    • Ask a peer manager or your own manager to sanity-check the list without the names first, if that's workable. A number that looks reasonable with context can look very different without it.
  5. Document the reasoning, not just the number

    The rating is the output. The reasoning is what actually gets tested in a calibration room, and it's also what forces sloppy logic into the open before anyone else sees it.

    • Write two or three sentences of reasoning per rating, referencing the actual evidence, before you have to defend it out loud.
    • If you can't write a clear reason for a number, that is itself useful information. It usually means the number isn't ready yet.
    • Keep the reasoning specific enough that someone with no context on the person could roughly follow why you landed where you did.
  6. Hold the line under pushback, but don't dig in for its own sake

    Calibration is where a rating actually gets tested against other managers' views. The goal isn't to defend your first draft at all costs. It's to know the difference between pushback that's revealing something you missed and pushback that's just disagreement.

    • If someone challenges a rating, go back to the evidence first, not to how confident you feel about your original read.
    • Being willing to move a number when the evidence genuinely supports it is not the same as caving to whoever argues loudest. The two can look similar from the outside, so know which one you're doing.
    • If you change a rating in the room, update your own reasoning notes to match. A mismatch between what's written and what was decided is exactly the kind of drift this whole process exists to catch.

That's the whole playbook

None of this replaces judgement. It is the checklist for catching the moments where judgement quietly turns into habit, or into bias.

Work through this against your own team

The Perform Review AI Manager Coach does two things this page cannot, because both are built from your own logged wins and your own reviews.

Guides: When a review of a direct report is written, your AI Coach can turn it into preparation for the meeting itself: talking points anchored to your strongest evidence, the honest gaps between their case and your rating, the pushback you are likely to get with answers to it, and a script you can say out loud.

Monthly Sessions: A read on your team every month, person by person: who is thriving, who is under-evidenced before calibration, the patterns across the team, and the conversations worth having this month. Reply to any part of it and the AI Coach answers on that specific point.

Both are included with and exclusive to Pro. See Pro.

More Manager Guides