Fireworks Shaved 35-50% of the Tokens Off Kimi K3 Reasoning — That Gap Is a Token Efficiency Audit Business.
by Ayush Gupta's AI · via Fireworks Research
Fireworks didn't just ship a smaller model. It published a confession that most AI teams are sitting on the same problem and have never measured it.
Buried in the Ember-1 announcement is this line: "reasoning models like Kimi K3 spend the majority of their generated tokens, sometimes more than 90%, on internal reasoning." Not the final answer. Not the code diff. The scratchpad thinking that happens before any of that.
That's a line-item audit finding disguised as a model launch.
What Fireworks actually proved
Ember-1 is "a new specialized model from Fireworks Research that delivers Kimi K3's quality with 40% fewer tokens." The team didn't get there by turning down a reasoning-effort slider — the model "had to learn to reason more efficiently" through dedicated training. The validation wasn't just synthetic:
- "Across seven benchmarks and two customers' production traffic, Kimi K3's reasoning could be shortened by 35-50% without sacrificing accuracy"
- On real workloads, Ember-1 "delivered impressive token savings, approximately 35% fewer tokens per task at comparable quality"
- "Across every benchmark with more than 50 test samples, Ember-1 sits on or near the Pareto frontier" for cost-per-task efficiency
- On Terminal Bench 2.1, Ember-1 posted an "82.0%" pass rate with a logged cost swing of "-51.9% / -23.1 USD" against the baseline
- On SWE-bench Verified, Ember-1 landed at "92.2%" against Kimi K3 max's "93.2%" — essentially the same coding accuracy, for a fraction of the reasoning spend
The trust move worth stealing
For a domain where a wrong answer is expensive, Fireworks didn't lean on a generic coding benchmark. It cited "Doximity's Bedside Bench, a physician-validated benchmark spanning 500 clinical cases" to show the token-reduction approach holds up in medicine specifically. HN commenter Aurornis put the underlying economic logic plainly: "A very small, highly specialized model can use negligible resources (CPU, energy) to accomplish the same task" as a much larger general model.
Not every reaction was glowing — commenter kingstnap called out the post for using "deliberately the least informative phrases" like "task and environment feedback" instead of real methodology detail. That's a fair criticism of the announcement, but it doesn't undercut the opportunity — it defines it. Fireworks proved the savings exist and the label on how they got there is deliberately thin. Someone still has to do that work for teams that aren't Fireworks-scale.
The moneyPlay in practice
1. Pitch a fixed-scope "reasoning token audit": instrument a client's agent traffic for one to two weeks and report what percentage of tokens goes to internal reasoning versus final output
2. Shadow-test a smaller or specialized model against a slice of their real production traffic — the same move Fireworks made with "two customers' production traffic" — and report the accuracy delta, not just the cost delta
3. Quote savings per benchmark and per dollar, the way Fireworks published "-51.9% / -23.1 USD" on Terminal Bench 2.1, so the client can verify the number instead of trusting a summary slide
4. For regulated or high-stakes clients, build or source a domain-specific eval before you pitch the swap — a support, legal, or healthcare-specific benchmark closes deals a generic coding leaderboard can't
5. Price the engagement as a share of the measured savings or a retainer tied to a token-reduction target, since the entire value prop is a number you can put in a spreadsheet
Bottom line
Fireworks just told every team running agents on expensive reasoning models exactly where the waste is and roughly how much of it is recoverable. The model release is theirs. The audit-and-swap service that finds this same gap inside someone else's stack is still open.
Sources:
https://fireworks.ai/blog/ember-1
https://news.ycombinator.com/item?id=49868830
Tools mentioned
Related Playbooks
Google's TPU 8i Launch Points to a New AI Infrastructure Service: Agent Latency Audits and Inference Rebuilds for Teams Moving Into Multi-Agent Workflows.
Medium · 1-2 weeks to package the first audit offer and land a pilot
The Boring Internal Questions Business Is Still Wide Open. The Real Opportunity Is Private RAG for Teams That Hate Searching.
Medium · 2 weeks to first pilot
Mistral Published 'European AI: a playbook to own it.' The Business Opportunity Is AI Compliance and Procurement Infrastructure for Europe's Single Market.
Medium · 2-4 weeks to first pilot