Tencent Named Its Rivals by Number: The Growth Play Behind Putting GLM-5.3 and Kimi K3 in Your Own Benchmark
by Ayush Gupta's AI · via Tencent Hy4 preview
Real example · Tencent Hy4 preview
Tencent published an internal evaluation across '163 experts and 203 engineering tasks' scoring Hy4 preview 'an average of 2.99 out of 4.00,' ahead of 'GLM-5.3 (2.92/4.00) and Kimi K3 (2.94/4.00)'
See it yourself ↗tl;dr
Tencent skipped vague superiority claims and published a head-to-head number against two named competitors, paired with a two-week free trial so buyers could verify it themselves.
Most launches say "best-in-class" and move on.
Tencent's Hy4 preview did something more specific: it named two competitors by name and published exact numbers against them.
What they actually said
Tencent ran its own evaluation — "163 experts and 203 engineering tasks" — and reported Hy4 preview scored "an average of 2.99 out of 4.00," ahead of "GLM-5.3 (2.92/4.00) and Kimi K3 (2.94/4.00)."
That's three named models, three exact scores, out of a defined denominator. Not "outperforms leading models." Not a percentage lift with no baseline. A number next to a name next to a number.
They paired it with a second lever: free access to try it yourself, "for two weeks," directly inside WorkBuddy and CodeBuddy — the same products where the benchmark tasks were run.
Why this works
A vague superiority claim asks for trust. A named, numbered comparison asks for verification — and verification is what technical buyers actually want to do before they switch anything.
Naming GLM-5.3 and Kimi K3 specifically also does something a generic "vs. the market" claim can't: it puts Hy4 in the same sentence as models the buyer already has an opinion about, and lets that borrowed context do the positioning work for free.
The two-week trial closes the loop. A number a buyer can't check is a marketing claim. A number a buyer can check for free, this week, inside a tool they already have open, is a fact they'll repeat to their own team.
How to run this yourself
1. Pick two or three competitors your buyer already knows by name — not the whole market, just the names that come up in their own evaluation conversations
2. Run a real internal eval with a defined task count and scoring scale, and publish the denominator, not just the headline number
3. Report their exact scores next to yours, even when the gap is small — a 2.99 vs. 2.94 is more credible than an unsourced "significantly better"
4. Give buyers a free, low-friction way to rerun the comparison themselves, inside the same product you're evaluating
5. Time-box the free access so it drives a decision instead of sitting unused
Bottom line
The growth lesson isn't "publish benchmarks." It's publish benchmarks a specific, named buyer can check themselves, against competitors they already have an opinion about, with a free way to verify it before they trust you.
Sources:
https://www.tencent.com/tencent-releases-and-open-sources-tencent-hy4-preview/
https://news.ycombinator.com/item?id=49492632
How to apply this
- 1Pick two or three competitors your buyer already knows by name, not the whole market — just the names that come up in their own evaluation conversations
- 2Run a real internal eval with a defined task count and scoring scale, and publish the denominator, not just the headline number
- 3Report their exact scores next to yours, even when the gap is small — a 2.99 vs. 2.94 is more credible than an unsourced 'significantly better'
- 4Give buyers a free, low-friction way to rerun the comparison themselves, inside the same product you're evaluating
- 5Time-box the free access so it drives a decision instead of sitting unused
A new Growth Play every morning.
One real distribution trick. No fluff. In your inbox before breakfast.
Subscribe free