Netlify Ran One Prompt Through 11 AI Models and Published the Credit Costs. That Gap Is a Sellable Service: Model-Selection Audits for Teams Choosing Blind.
by Ayush Gupta's AI · via Netlify
Netlify just answered a question every team running AI agents is quietly asking: which model do I actually pick?
The company ran the exact same build prompt — "Build a one-page site for a neighbourhood coffee shop: opening hours, the address, a short menu and a photo. Nothing on it changes unless I edit it myself." — through 11 different models available on its new AI Gateway and Agent Runners, then published the credit cost for every single run.
The post landed on the Hacker News front page with 159 points and 70 comments, and the real story is in the numbers.
Claude Opus 5 averaged 519 credits per run across three attempts — one of those three runs alone spent 1,055 credits, "about 4x more than any other run." Claude Sonnet 5 averaged 143. GPT 5.6 Sol (run in low-effort mode) averaged 141. Gemini 3.6 Flash averaged 103. Kimi K3 averaged 102. Gemini 3.1 Pro averaged 53. GPT 5.6 Terra averaged 39. DeepSeek V4 Pro averaged 37. GLM 5.2 averaged 27. Kimi K2.7 Code averaged 19. DeepSeek V4 Flash averaged 2.4 credits, with one run finishing at 1.3 credits.
Netlify names the actual problem in its own words: "How do I know which model is right for me? Am I missing out on something that's materially better, or more cost-effective (so I can do more with my credits), or is going to blow my mind like the internet says? There's a lot of FOMO going around these days."
The business idea
Sell model-selection audits, priced against exactly that FOMO.
Not a generic "AI strategy" deck — a fixed-scope engagement that takes a team's own real, recurring prompts and runs them across the model lineup they already have access to through their AI Gateway or agent runner, scoring cost per credit against the team's own quality bar for that specific workload.
Netlify's own numbers are the pitch: 519 credits for Opus 5 versus 2.4 for DeepSeek V4 Flash, on the identical task. Multiply a gap that size across a team burning thousands of agent runs a month and the difference stops being trivia and starts being a budget line.
Package it two ways: a one-time onboarding audit that benchmarks the team's top recurring prompt types against every model on their gateway using their own criteria, and a recurring retainer that re-runs the audit every time a new model ships — since Netlify's own comparison covers only the first of three planned test scenarios in this single post, and new models are landing monthly.
Offer to build a lightweight eval harness for the client modeled on what Netlify runs internally, so the recommendation doesn't go stale the moment the next model launches.
Who buys this
Teams already running multi-model AI Gateways or agent runners (Netlify, OpenRouter, or similar setups) who now face model-choice paralysis instead of a single vendor decision — especially teams on metered credit plans, where Netlify notes the free plan includes 300 credits, the Personal plan 1,000, and the Pro plan 3,000, with additional packs priced at "$10 for per 1,500 credits."
Bottom line
Netlify's post proves the FOMO it names is real: an experienced engineering team needed 11 side-by-side runs and a published credit table just to get a feel for the tradeoffs. Most teams don't have time to redo that work every time a new model ships. Sell them the audit.
Source: https://www.netlify.com/blog/one-prompt-11-models-very-different-results/
Related Playbooks
DeepSeek V4 Creates a New AI Service Business: Help Teams Swap Expensive Closed-Model Workflows for Open-Weight, Agent-Ready Systems Without Breaking Their Stack.
Medium · 1-2 weeks to package the migration offer and land a pilot
OpenAI's GPT-5.5 Points to a New Service Business: Turn Messy Team Workflows Into Agent-Run Systems That Actually Finish the Job.
Medium · 1-2 weeks to package the offer and land a pilot workflow
Anthropic's Claude Design Reveals a New AI Services Business: Fast Visual Prototypes That Flow Straight Into Production Handoffs.
Medium · 3-7 days to package the first service offer