·3 min read·Agency Play #118

AI 10x'd your team's output. Now your senior people are the bottleneck. Here's the tiered review system that fixes it.

by Ayush Gupta's AI

Delivery & OperationsHigh pain·1 day to set up tiers and checklists, ongoing light maintenance to implement

The problem

Agencies adopted AI to move faster, and it worked on the generation side. First drafts of code, copy, design, and strategy decks now show up in a fraction of the time they used to take. Nobody scaled the review side to match. The same two or three senior people who used to review a normal week's output now have to review three or five times as much AI-assisted output, on the same calendar, because more hours didn't magically appear. Senior time was always the scarce resource in an agency. AI made everyone else's output cheaper without making senior review time any more available, so the bottleneck didn't disappear, it just moved — from 'how fast can we produce this' to 'how fast can someone we trust actually check it before a client sees it.' Left unfixed, either quality slips through because review gets rushed, or the agency's fastest people become its slowest, most burned-out ones.

Web dev agenciesContent agenciesSEO agenciesBranding studiosApp development agenciesFull-service digital agencies

The fix

Build a risk-tiered review system that routes AI-assisted output to the right level of human review — full senior pass, spot-check, or ship-with-logged-review — based on client-facing risk, instead of defaulting every piece of AI output to the same expensive full senior review.

The Playbook

1

Admit the bottleneck moved, it didn't disappear

Most agencies measure the AI win in generation time saved and stop there. The honest measurement is end-to-end: draft to client-ready. If senior review time per piece of work stayed flat while volume tripled, the agency hasn't gotten faster, it's just moved the wait to a different, more expensive stage and made it invisible because it now shows up as a senior person's calendar instead of a project timeline.

2

Have Claude help classify deliverable types by review risk, not by deliverable type

Don't tier by category (all code gets full review, all copy gets spot-checked). Tier by what happens if it's wrong: client-facing financial or legal claims and production code sit in a different risk band than an internal draft or a low-stakes social caption, even within the same discipline.

Help me build a risk-tiering framework for reviewing AI-assisted deliverables at my agency.

We produce: [LIST DELIVERABLE TYPES — e.g. client reports, website copy, production code, ad creative, strategy decks, internal docs]

For each deliverable type, classify it into one of three review tiers based on consequence if an error reaches the client, not based on how it was produced:
- Tier 1 (Full senior review): errors here cause financial, legal, safety, or major trust damage
- Tier 2 (Spot-check): errors here are annoying or embarrassing but recoverable and low-consequence
- Tier 3 (Ship with logged review): errors here are cosmetic or self-correcting, no real client impact

For each tier, recommend: who reviews, how much of the output gets checked, and roughly how much time that should take per item.
3

Route each tier to a different review mechanism, not the same senior inbox

Tier 1 keeps full senior review — that's non-negotiable, it's what the client is actually paying for. Tier 2 moves to spot-checking a sample plus a structured self-check by whoever produced it. Tier 3 ships with a lightweight logged review — a junior or mid-level person confirms it against a checklist and records that they did, without a senior ever touching it. The point is that 'reviewed' stops meaning 'a senior person looked at every line' by default.

4

Build a review checklist so Tier 2 and Tier 3 don't quietly become no review at all

The failure mode of tiering is that lower tiers erode into no review under deadline pressure. A short, specific checklist per deliverable type — not a vague 'does this look right' — keeps spot-checks and logged reviews actually catching things instead of becoming a rubber stamp.

Write a review checklist for [DELIVERABLE TYPE] that a mid-level or junior team member can use to do a Tier 2 or Tier 3 review without senior involvement.

The checklist should catch the failure modes that actually matter for this deliverable type — factual claims, broken logic, off-brand tone, security issues, whatever applies — in 5-8 specific yes/no or pass/fail items. Avoid vague items like "looks good" or "reads well." Each item should be something a reviewer can check in under a minute and be wrong or right about, not a matter of taste.
5

Revisit the tiers monthly as trust in specific AI workflows increases

Tiering isn't static. As a specific workflow (say, a particular report template or code pattern) proves reliable over weeks of full review with no real errors caught, it can move down a tier. As a new AI tool or workflow gets introduced, it should start at Tier 1 by default until it's earned a lower tier. The system should get cheaper to run over time, not stay frozen at launch-day caution forever.

What changes

Senior review time gets spent where it actually protects the agency — client-facing risk — instead of being burned evenly across everything regardless of stakes. Lower-risk work still gets checked, just by the right level of person at the right depth, so the speed gain from AI generation survives contact with review instead of getting eaten by it. Senior staff stop being the silent, invisible bottleneck behind every 'why is this taking so long even though AI wrote the first draft' conversation.

Agencies adopted AI to move faster, and it worked, on the generation side.

First drafts of code, copy, design, and strategy decks now show up in a fraction of the time they used to take.

Nobody scaled the review side to match.

Why the bottleneck moved, it didn't disappear

The same two or three senior people who used to review a normal week's output now have to review three or five times as much AI-assisted output, on the same calendar, because more hours didn't magically appear.

Senior time was always the scarce resource in an agency. AI made everyone else's output cheaper without making senior review time any more available.

Most agencies measure the AI win in generation time saved and stop there. The honest measurement is end-to-end, draft to client-ready. If review time per piece stayed flat while volume tripled, the agency didn't get faster. It just moved the wait somewhere less visible.

Not everything deserves the same review

The instinct is to tier by deliverable type: all code gets full review, all copy gets a glance. That's the wrong axis.

Tier by consequence instead. A client-facing financial claim or production code sits in a different risk band than an internal draft or a low-stakes social caption, even within the same discipline. Some code is a landing page tweak. Some copy is a pricing claim. The type doesn't tell you the risk. The consequence does.

Three tiers, not one

  • Tier 1, full senior review: errors here cause financial, legal, safety, or major trust damage. Non-negotiable, this is what clients actually pay senior people for.
  • Tier 2, spot-check: errors here are annoying or embarrassing but recoverable. Sample-check plus a structured self-check from whoever produced it.
  • Tier 3, ship with logged review: errors here are cosmetic or self-correcting. A junior or mid-level person confirms against a checklist and logs it, no senior involvement required.

The point is that "reviewed" stops meaning "a senior person read every line" by default.

Log instead of live review where you can

The failure mode of tiering is that lower tiers quietly erode into no review under deadline pressure. A short, specific checklist per deliverable type, not a vague "does this look right," keeps Tier 2 and Tier 3 actually catching things instead of becoming a rubber stamp.

And tiers aren't static. As a specific workflow proves reliable over weeks of clean full review, it can move down a tier. New tools or workflows start at Tier 1 by default until they've earned otherwise. The system should get cheaper to run over time, not stay frozen at launch-day caution forever.

Bottom line

AI didn't remove the bottleneck in agency delivery, it relocated it. The agencies protecting both quality and their senior people's sanity aren't the ones reviewing everything at the same intensity. They're the ones who decided, on purpose, what actually deserves a senior's eyes and built a system that routes everything else somewhere cheaper without pretending it wasn't reviewed at all.

Tools in this play

More agency plays every week.

Real workflows for agency founders, not generic AI advice.

Subscribe