Cognition's SWE-2 Launch Shows the Move: Sell the Efficiency Delta Against Your Own Last Version, Not Just the Win Rate
by Ayush Gupta's AI · via Cognition / SWE-2
Real example · Cognition / SWE-2
Led its launch with operational efficiency numbers against its own prior model — '58% fewer turns and costing 81% less on average' and 'a median of 18 steps, compared with 48 for SWE-1.7' — alongside accuracy claims against public benchmarks
See it yourself ↗tl;dr
Cognition had two kinds of numbers to lead SWE-2's launch with: accuracy percentages against public benchmarks, or efficiency numbers against its own predecessor. The efficiency numbers went largely uncontested in a 131-comment HN thread; the accuracy numbers got called 'benchmaxxing.'
The Play
Cognition had two kinds of numbers to lead its SWE-2 launch with: accuracy percentages against public benchmarks, or operational efficiency numbers against its own prior model. It led hardest with the second — and that's the number Hacker News's 131-comment thread mostly left alone.
Why this matters
A buyer deciding whether to upgrade an AI tool they already pay for doesn't need to be convinced the new version is smarter in the abstract — they need to know what it costs to keep doing what they're already doing. "Fewer turns, fewer steps, lower cost" answers that question directly, in units the buyer already tracks. "Higher score on a benchmark you've never run" doesn't, and it invites the exact scrutiny SWE-2's accuracy claims got: a commenter went and found SWE-2's Terminal-Bench 4 score (27.3%) sitting below two competitors' independently reported numbers on the same suite (DeepSeek V4.1 Flash at 31.2%, Sol at 37.3%).
The asymmetry is the lesson. Self-referential efficiency numbers — this version versus your own last version — are hard to argue with because there's no competing narrative to check them against. Cross-vendor accuracy numbers are trivially checkable against a dozen other vendors' launch blogs, and technical crowds will do exactly that.
How to run this play
1. Lead your comparison against your own previous version's numbers before you lead with any competitor comparison — a same-product delta can't be contradicted by someone else's benchmark run
2. Report the unit your buyer already has a budget line for — turns, steps, dollars, minutes — instead of a benchmark percentage they have no independent way to sanity-check
3. Pair every efficiency figure with the exact number it replaced ("18 steps, compared with 48") so the improvement is legible without the reader doing arithmetic
4. Treat any accuracy or win-rate claim as the part of your launch that will get independently cross-checked hardest — publish it, but don't lean your headline on it
5. Publish your full benchmark table, including the categories where your number is weakest — a reader who finds the gap themselves treats you as evasive; a reader who sees you disclosed it treats the rest of your numbers as credible
6. Keep your efficiency headline short enough to survive being repeated third-hand — "58% fewer turns, 81% less cost" fits in a comment reply; a benchmark methodology paragraph doesn't
Bottom line
The number that spread from SWE-2's launch wasn't a win-rate percentage — it was 58% fewer turns and 81% less cost against its own predecessor. The accuracy claims sitting right next to it got picked apart as "benchmaxxing" within the same thread. Buyers trust a company's own before/after math more than they trust its scorecard against everyone else.
Sources:
https://cognition.com/blog/swe-2
https://news.ycombinator.com/item?id=49645443
How to apply this
- 1Lead your comparison against your own previous version's numbers before you lead with any competitor comparison — a same-product delta can't be contradicted by someone else's benchmark run
- 2Report the unit your buyer already has a budget line for — turns, steps, dollars, minutes — instead of a benchmark percentage they have no independent way to sanity-check
- 3Pair every efficiency figure with the exact number it replaced ('18 steps, compared with 48') so the improvement is legible without the reader doing arithmetic
- 4Treat any accuracy or win-rate claim as the part of your launch that will get independently cross-checked hardest — publish it, but don't lean your headline on it
- 5Publish your full benchmark table, including the categories where your number is weakest — a reader who finds the gap themselves treats you as evasive; a reader who sees you disclosed it treats the rest of your numbers as credible
- 6Keep your efficiency headline short enough to survive being repeated third-hand — '58% fewer turns, 81% less cost' fits in a comment reply; a benchmark methodology paragraph doesn't
A new Growth Play every morning.
One real distribution trick. No fluff. In your inbox before breakfast.
Subscribe free