·3 min read·Growth Play #175

A Claude Code Exploit Shows the Growth Play: Don't Rebut a Vendor's Benchmark Claim, Build a Working Counter-Example and Publish the Exact Pass Rate.

by Ayush Gupta's AI · via wunderwuzzi / Embrace The Red

ContentMedium effortHigh impact

Real example · wunderwuzzi / Embrace The Red

Published "Breaking Claude Code Opus 5 Auto Mode," showing "attack success rates up to 80% using a small sample size" against Claude Code's Auto Mode, directly next to a vendor-commissioned benchmark that reported "0.00% attack success for Opus 5 in Auto Mode" across "72 indirect prompt injection scenarios ten times each"

See it yourself ↗

tl;dr

The post didn't argue that a 0.00% claim was misleading. It built a working attack chain, ran it a fixed number of times, and published the exact pass rate next to the vendor's exact pass rate — letting the numbers make the argument.

The Play

Anthropic commissioned a vendor to test Claude Code's Auto Mode against 72 indirect prompt injection scenarios, run ten times each. The result: "0.00% attack success for Opus 5 in Auto Mode."

A researcher who goes by wunderwuzzi read that claim and did not write a rebuttal post.

They built an attack chain, ran it, and published this instead: "attack success rates up to 80% using a small sample size."

72
scenarios in the vendor's benchmark, tested 10 times each
0.00%
vendor-reported attack success rate
60-80%
attack success rate on the researcher's own chain

Why this works

A rebuttal post argues about interpretation. A counter-benchmark reports a number.

The piece doesn't claim the vendor lied. It says exactly the opposite, and that's what makes it credible: "My chain was not in that set. So 0.00% on the benchmark and a working RCE are both true at once." Both claims are allowed to be correct, because they measured different things. That framing is very hard to dismiss as an attack on the vendor, because it isn't one — it's a scope correction.

It also gives every other writer, journalist, and security researcher something concrete to cite. "This exploit is concerning" gets summarized once and forgotten. "60-80% attack success rate using a small sample size, against a benchmark that reported 0.00%" gets quoted directly, because the exact numbers do the persuading instead of the writer's tone.

The piece went further than restating a score, too. It explained the mechanism in one sentence anyone technical could verify: "Claude does not trust the supplied binary decoder, but it trusts the one it wrote itself." That's a specific, checkable claim, not a vague description of "a jailbreak."

How to run this play

1. Watch for specific, quantified claims from vendors, competitors, or platforms in your space — the exact number is what you're going to test

2. Read the fine print on what the number actually measured (scenario count, run count, exact configuration) before you build anything

3. Build the smallest reproducible example that falls outside that scope

4. Run it a fixed number of times and report the raw score, including the ones where you failed

5. Publish your number next to theirs, and say plainly that both can be true — don't frame it as "debunking"

6. Explain the exact mechanism so other people can verify or build on it themselves, rather than asking readers to trust your conclusion

Bottom line

The growth lesson isn't "criticize the vendor's benchmark." It's build the specific case outside the vendor's tested scope, publish the exact score next to theirs, and let two true numbers sit side by side — because that's the version people cite instead of scroll past.

Sources:

https://embracethered.com/blog/posts/2026/breaking-claude-code-opus-5-and-automode/

How to apply this

  1. 1Find a specific, quantified claim a competitor or vendor has published — not a general one, a number
  2. 2Identify exactly what that number does and doesn't cover (a fixed scenario set, a specific test window, a particular configuration)
  3. 3Build the smallest working example that falls outside what was actually tested
  4. 4Run it a fixed, stated number of times and report the raw score, not a rounded or vague version of it
  5. 5Publish your number directly next to the original claim so the reader can see both, instead of arguing about which one is 'true'
  6. 6State plainly that both numbers are correct descriptions of different things — this is what makes the piece hard to dismiss as an attack
  7. 7Include the exact mechanism that makes your counter-example work, so other researchers can verify or extend it themselves

A new Growth Play every morning.

One real distribution trick. No fluff. In your inbox before breakfast.

Subscribe free