Cognition's SWE-2 Scored 92.8% and 27.3% on the Same Benchmark Family. That Gap Is a New Coding-Agent Evaluation Business.
by Ayush Gupta's AI · via Cognition
Cognition's SWE-2 launch reads like a routine model release until you notice two conflicting numbers sitting in the same blog post — and how differently a 131-comment Hacker News thread treated each one.
What actually happened
Cognition post-trained SWE-2 from Kimi K3, a 2.8-trillion-parameter base model, calling it the first time RL had been "scaled...to the multi-trillion-parameter regime." The launch blog leaned on two kinds of numbers: accuracy gains (50.0% vs 42.0% on FrontierCode 1.1 Main, 73.0% vs 37.7% on DeepSWE 1.1, 92.8% vs 81.5% on Terminal-Bench 2.1) and operational gains against its own predecessor (fewer turns, fewer steps, lower cost). Cognition also published a trustworthiness eval showing SWE-2 "passed 98.0% of attempts overall: 99.8% in English, 95.2% in Simplified Chinese, and 99.1% in Traditional Chinese." The model shipped inside Devin Desktop, CLI, Web, and Fusion.
The HN thread split along the seam in that data. Commenters flagged the gap between Terminal-Bench 2.1 (92.8%) and Terminal-Bench 4 (27.3%) as evidence of "benchmaxxing" — overfitting to a suite the model has effectively seen before rather than genuine capability on a fresh one. One commenter went further and cross-referenced that 27.3% against independently reported Terminal-Bench 4 scores for DeepSeek V4.1 Flash (31.2%) and Sol (37.3%). A separate, recurring complaint had nothing to do with accuracy: "Why would I use this over DeepSeek Flash 4.1?" given open-weight models are "nearly as cheap on API usage rates," plus a lock-in objection — "I'm disappointed to see I need to use a bespoke platform to interact with this agent" — that echoed Cognition's earlier reputation problems with Devin's autonomous-Upwork demo going "off the rails."
What this exposes
- Public benchmarks disagree with each other by construction — the same model scored 92.8% and 27.3% on two versions of the same benchmark family, a 65-point swing that has nothing to do with the model changing and everything to do with which suite got quoted
- Vendors get to pick which comparison to publish — SWE-2's launch table compares itself to its own predecessor and to Fable 5.1, not to the independent Terminal-Bench 4 scores a commenter had to go find manually
- Buyers have no standing benchmark for their own codebase — every number in the debate came from a public suite that measures something adjacent to, not identical to, any given team's actual backlog
- Platform lock-in is a live objection, not a hypothetical — teams evaluating a coding agent are weighing "bespoke platform" switching cost alongside raw capability, and that number doesn't show up in any vendor's benchmark table
The business idea
Any team about to standardize on a coding agent — Devin, DeepSeek-based tooling, Sol, or something newer next quarter — is choosing between benchmark tables built to make each vendor look best on the axis where it wins:
- Pull a real slice of the client's own backlog and run it through 3-4 current leading agents under identical conditions
- Score both accuracy and the operational metrics Cognition itself highlighted as the real story — turns per task, steps to first edit, cost per resolved ticket — since those determine ongoing spend, not a one-time win rate
- Cross-check any vendor-supplied benchmark number against at least one independent source before it goes into the recommendation
- Price out platform lock-in cost as its own line item in the audit report
- Deliver a short decision memo naming a primary model plus a fallback, since the ranking is likely to flip again within a quarter
Why this works now
New coding-agent releases are arriving roughly monthly, each with a benchmark table built to spotlight its own win. Teams don't have the bandwidth to independently verify every claim, and HN just demonstrated in public what happens when someone does the cross-checking work: a 65-point discrepancy across the same model on two versions of one benchmark family. That verification labor is exactly what a fixed-scope audit sells.
Bottom line
SWE-2's launch blog and its own comment section told two different stories from the same numbers — one built on the vendor's chosen comparisons, one built on a reader's five minutes of cross-referencing. The gap between those two stories is the audit.
Sources:
https://cognition.com/blog/swe-2
https://news.ycombinator.com/item?id=49645443
Tools mentioned
Related Playbooks
DeepSeek V4 Creates a New AI Service Business: Help Teams Swap Expensive Closed-Model Workflows for Open-Weight, Agent-Ready Systems Without Breaking Their Stack.
Medium · 1-2 weeks to package the migration offer and land a pilot
OpenAI's GPT-5.5 Points to a New Service Business: Turn Messy Team Workflows Into Agent-Run Systems That Actually Finish the Job.
Medium · 1-2 weeks to package the offer and land a pilot workflow
Anthropic's Claude Design Reveals a New AI Services Business: Fast Visual Prototypes That Flow Straight Into Production Handoffs.
Medium · 3-7 days to package the first service offer