Your best prompter looks like your best performer on paper. Here's the calibration system that tells them apart.
by Ayush Gupta's AI
The problem
AI collapsed the gap between your fastest and slowest employees' raw output, which means the volume-based signals most agencies still use in performance reviews — decks shipped, tickets closed, drafts turned around — no longer track skill the way they used to. The person clearing the most work might just be the heaviest AI user with the loosest editorial bar. The person clearing the least might be catching problems nobody else sees before they hit the client. Review cycles built around throughput are quietly promoting the wrong people and undervaluing the ones actually protecting the agency's quality — and almost nobody has updated the rubric to notice.
The fix
Replace volume-based review criteria with a calibration process that scores the judgment calls behind the work — what someone caught, what they escalated, what they chose not to ship — using Claude to structure the evidence so managers compare decisions instead of deliverable counts.
The Playbook
Name what your current review actually measures
Pull the last two review cycles for your team and look at what actually drove the ratings — deliverables shipped, hours logged, tickets closed, campaigns launched. Almost every agency review, even ones that claim to be qualitative, ends up anchored on a volume proxy because it's the easiest thing to point to in the moment. Write down, honestly, which of your current criteria are throughput in disguise. That list is what needs to change, not the whole review process.
Have Claude turn a work sample into a judgment-evidence brief
For each person being reviewed, pull three to five real deliverables from the cycle — the actual output plus whatever context exists on how it got there (Slack threads, revision history, client feedback). Feed it to Claude and have it extract the decisions, not just describe the output.
You are helping me prepare a performance review that evaluates judgment, not output volume.
Here is a work sample from this review cycle: [PASTE DELIVERABLE + SURROUNDING CONTEXT — Slack threads, revision history, client feedback, brief]
From this, extract:
1. What decisions did this person actually make — what to include, what to cut, what to flag, what to push back on
2. Where did they catch something before it became a problem (or fail to)
3. Where did they defer to AI output without visible scrutiny, if you can tell
4. What would a strong version of this decision have looked like vs a weak one, and which is closer to what happened
Don't summarize the deliverable itself — I can read that. Focus only on the judgment calls behind it. Flag anywhere you don't have enough context to assess a decision instead of guessing.Score decisions on a rubric, not deliverables on a count
Build the review rubric around a small set of judgment categories that apply regardless of role or output volume: what they caught, what they escalated correctly, what they shipped that they shouldn't have, and what they held back that they should have shipped. Score every person against the same four categories using the judgment-evidence briefs from step 2, not a tally of how much they produced.
Run a calibration pass across managers before ratings go final
The same judgment-evidence brief can read very differently depending on which manager reviews it, especially once volume stops being the anchor. Put every draft rating in front of a second manager or a small calibration group before it's finalized, and have Claude flag the gaps.
Compare these draft performance ratings against the judgment-evidence briefs they're based on.
Ratings and briefs: [PASTE DRAFT RATINGS + THE JUDGMENT-EVIDENCE BRIEFS THEY WERE BASED ON]
Flag:
1. Any rating that seems to track output volume more than the judgment evidence actually supports
2. Any two people with similar judgment evidence who received meaningfully different ratings
3. Any rating where the manager's stated reasoning doesn't match what's in the brief
I want inconsistencies surfaced before these ratings go to the team, not after.Tell the team the rubric changed, and why, before the next cycle
A rubric change that shows up unannounced in someone's review reads as arbitrary, even when it's an improvement. Explain the shift explicitly at the start of the next cycle: output volume stopped being a reliable signal once everyone had AI leverage, so reviews are now weighted toward the decisions behind the work. Give the team the same four judgment categories in advance so nobody is graded against a standard they never saw.
What changes
Reviews start rewarding the people actually protecting quality and catching problems, instead of quietly promoting whoever produces the most AI-assisted volume with the least scrutiny. Managers get a rubric that still works as AI capability keeps shifting, because it's anchored on judgment rather than a throughput number that AI can inflate for anyone.
AI gave your slowest employee and your fastest employee the same leverage. The gap in what they ship narrowed. The gap in whether what they ship is any good didn't. Most agency performance reviews haven't caught up to that, and they're still scoring the wrong half of the gap.
The real problem
For years, output volume was a decent enough proxy for skill in a review, because producing more, faster, usually meant someone was genuinely good. AI broke that correlation. Now the person clearing the most decks, drafts, or tickets in a cycle might just be running everything through AI with a thin editorial pass, while the person clearing the least might be the one actually reading every line, catching the thing that would've embarrassed the agency in front of a client, and choosing not to ship the first draft AI handed them.
Reviews built on the old proxy don't notice the difference. They see throughput, and throughput still looks like the easiest thing to reward — a clean number, easy to defend, easy to compare across a team. So agencies keep rating on it, quietly promoting the heaviest AI users with the loosest bar and undervaluing the people whose real contribution is judgment that never shows up in a deliverable count.
The fix
Stop scoring deliverables and start scoring the decisions behind them. Pull real work samples, extract what someone actually caught, escalated, shipped, or held back, and rate that against a small, consistent rubric — not a count of how much came out the other end. Calibrate ratings across managers before they go final, because judgment-based scoring is more subjective than a tally and needs a second read to stay fair. Then tell the team the rubric changed and why, so the new standard doesn't land as a surprise in someone's next review.
Why this matters
The agencies that keep reviewing on volume in an AI-leveraged world are running a slow-motion adverse selection process on their own team: rewarding confident AI users who ship fast and light, training everyone else that carefulness doesn't pay, and losing the people whose judgment is the actual reason clients trust the work. That's not a hypothetical risk sitting a few cycles out — it's already happening in every review that still leads with "how much did you ship this quarter."
Bottom line
AI made everyone faster. It didn't make everyone better, and your review process needs to be able to tell the two apart. Score the judgment, calibrate the ratings, and tell the team what changed — before the next cycle rewards the wrong person again.