Your client's open rates just cratered and nobody touched the copy. Here's the AI deliverability audit that catches the spam classifier before the client blames your work.
by Ayush Gupta's AI
The problem
Open rates on a client's email program drop 30-40% in a month, nothing about the strategy or list changed, and the first assumption everyone reaches for is bad subject lines or a tired offer. Increasingly, the real cause is upstream: Gmail and Outlook's spam classifiers have gotten aggressive about detecting the exact patterns AI-assisted email production creates — near-identical structure across sends, formulaic subject line formats, generic CTAs, and copy that reads as templated rather than human, even when it's technically well-written. The agency spends a week rewriting copy that was never the problem, while the client quietly loses trust in the channel.
The fix
Build a standing pre-send audit that checks both deliverability infrastructure (authentication, list hygiene, complaint rate) and AI-tell content patterns before a campaign ships, so a classifier problem gets caught and named before it shows up as a client complaint about performance.
The Playbook
Rule out the classifier before you touch the copy
Before rewriting anything, pull the actual diagnostic data: Google Postmaster Tools spam rate and IP/domain reputation, the ESP's own bounce and complaint reports, and an inbox placement test through a tool like GlockApps. If spam-folder placement jumped on the same send that open rates dropped, the problem is upstream of the copy, and no amount of subject-line polishing will fix it. Have Claude turn the raw exports into a one-page verdict so this doesn't turn into a week of guessing.
You're diagnosing an email deliverability drop for a client campaign.
Data: [PASTE POSTMASTER TOOLS EXPORT, ESP BOUNCE/COMPLAINT REPORT, INBOX PLACEMENT TEST RESULTS]
Context: open rate dropped from [X%] to [Y%] starting around [DATE]. No list changes, no send-time changes.
Give me:
1. Whether this looks like a reputation/classifier issue, a content-trigger issue, a list-quality issue, or a mix
2. The single strongest piece of evidence for your read
3. What would disprove your read, so I don't chase the wrong fix
One page. No hedging.Score the copy for the exact patterns the classifier is trained to catch
Gmail and Outlook's spam models are increasingly tuned to flag 'bulk sameness' — near-identical structure across sends, formulaic subject lines, generic CTAs, and copy with the telltale rhythm of AI-assisted drafting rather than human variation. Run every campaign through a checker before it ships, not after the report shows a drop, using the same signals a classifier would weight.
Score this email campaign for spam-classifier risk, specifically the "AI bulk sameness" signals Gmail/Outlook models weight heavily:
Email: [PASTE FULL EMAIL INCLUDING SUBJECT LINE]
Last 5 sends to this list for comparison: [PASTE OR SUMMARIZE]
Flag:
1. Structural repetition vs. the last 5 sends (same skeleton, same CTA placement, same sentence rhythm)
2. Formulaic subject line patterns (numbered lists, urgency framing, templated brackets)
3. Generic CTA language that reads as templated rather than specific to this send
4. Anything that reads as AI-drafted-and-shipped without a human editing pass
Rate overall risk Low/Medium/High and give the two highest-leverage fixes.Make the audit a pre-send gate, not a post-mortem
The only version of this that actually protects the client relationship runs before the campaign goes out, not after the performance report raises questions. Build it into the send checklist: authentication check, complaint-rate check, and the AI-tell content score all have to clear before a campaign leaves the queue. This is a five-minute gate once it's built, not a new review cycle.
Fix the infrastructure fundamentals, because they stack with content flags
A classifier doesn't weigh content in isolation — misaligned SPF/DKIM/DMARC, a stale or unengaged segment of the list, and a complaint rate above 0.3% all compound with formulaic copy to tank placement faster than any one factor alone. Run an authentication and list-hygiene pass on every client domain sending meaningful volume, and fix the infrastructure gaps before optimizing copy further. Google and Yahoo's bulk sender requirements are the floor here, not a nice-to-have.
Report the cause, not just the metric, before the client asks
When a client sees open rates drop, their first instinct is to blame the offer or the copy, because that's the part they can see. Get ahead of it: if the audit finds a classifier or infrastructure cause, say so in the next report, in plain language, with what's already been fixed. A client who hears 'we caught a deliverability issue and fixed it' trusts the agency more than one who watches a metric drop silently for a month.
Write a client-facing paragraph for a monthly report explaining a deliverability issue we caught and fixed.
Issue found: [E.G. "SPAM CLASSIFIER FLAGGING DUE TO REPETITIVE CAMPAIGN STRUCTURE" / "DKIM MISALIGNMENT" / "LIST COMPLAINT RATE ABOVE THRESHOLD"]
What we did about it: [FIX APPLIED]
Expected recovery timeline: [TIMEFRAME]
Tone: calm, plain-language, no jargon dump, positions us as the team that caught it before it became a bigger problem. 4-5 sentences.What changes
Deliverability drops get diagnosed correctly instead of triggering a week of copy rewrites that don't address the actual cause, campaigns get checked for classifier risk before they ship instead of after a client notices, and the agency reports deliverability issues proactively instead of getting caught flat-footed by a client who saw the drop first.
Every agency running email for a client has had this conversation: open rates dropped, nobody changed anything meaningful, and the client wants to know why the copy stopped working.
Most of the time, the copy didn't stop working. The inbox stopped trusting it.
Gmail and Outlook have spent the last two years making their spam classifiers more aggressive about a specific pattern: bulk sameness. Not spammy words, not blocklisted links — just the structural fingerprint of email that's produced at volume with a repeatable formula. Same skeleton every send. Same CTA placement. Same subject line pattern with the variable swapped out. That fingerprint used to just look like "efficient email marketing." Now it's exactly what a classifier trained to catch AI-assisted bulk production is looking for.
The agencies getting hurt by this aren't the ones doing anything wrong. They're the ones who got efficient at producing consistent, on-brand campaigns quickly — which, from a spam model's point of view, looks identical to a low-effort bulk operation.
The trap: diagnosing the wrong layer
When open rates drop, the instinct is to look at the thing you can see and control: the copy. So the team rewrites subject lines, tests new CTAs, maybe brings in a new writer. None of it moves the number, because the problem was never the copy quality. It was placement — the email never reliably reached the inbox to begin with.
Diagnose the layer before you touch the copy
The fix starts with ruling things out in the right order. Pull Google Postmaster Tools data, the ESP's own complaint and bounce reports, and run an actual inbox placement test. If spam-folder placement jumped on the same send where opens dropped, that's the answer, and it's upstream of anything a copywriter can fix.
Score for the exact signals the classifier weighs
Once infrastructure is ruled in or out, score the content itself for the "AI bulk sameness" signals: structural repetition across sends, formulaic subject line patterns, generic templated CTAs, and copy that reads as produced rather than written. This isn't about banning AI from the drafting process — it's about catching the tells before a classifier does, the same way you'd catch a typo before it ships.
Build it as a gate, not a post-mortem
The version of this that actually protects the client relationship runs before send, not after the report shows a drop. Authentication check, complaint-rate check, content-tell score — all have to clear before a campaign leaves the queue. Once it's built this takes minutes, not a new review cycle.
Fix the infrastructure fundamentals
Content flags stack with infrastructure gaps. Misaligned SPF/DKIM/DMARC, a stale list segment, or a complaint rate creeping above 0.3% will tank placement faster in combination than any single factor alone. Google and Yahoo's bulk sender requirements are the floor every client domain needs to clear, not an optional upgrade.
Report the cause before the client finds it
If the audit turns up a real issue, say so proactively, in the next report, in plain language. A client who hears "we caught a deliverability issue and already fixed it" trusts the agency more after the incident than before it. A client who watches a metric drop for a month with no explanation starts shopping for a new agency.
The honest caveat
This audit won't save a genuinely weak offer or a list that's been over-mailed into fatigue — those are real problems the classifier is also correctly flagging, and no amount of technical fixing will paper over content that deserves to underperform. The value here is narrower and more specific: separating "the email is bad" from "the email never reliably arrived," so the team spends its time fixing the actual problem instead of polishing copy that was never broken.
The classifiers are only getting better at this. The agencies that build the pre-send check now will spend less time firefighting deliverability surprises next quarter. The ones that keep diagnosing every open-rate drop as a copy problem will keep rewriting subject lines that were never the issue.