Castform Shows a New AI Service: Post-Train a Small Open Model on a Client's Own Data So Retrieval Costs Drop 100x Without Losing Accuracy.
by Ayush Gupta's AI · via Neon
Castform didn't launch a bigger model.
It launched a way to make a small one match a frontier model on one specific job, for a lot less money.
That's the business idea.
What actually happened
Neon's writeup on Castform states the result plainly: "A 4B open-source model post-trained with Castform retrieved search results as accurately as GPT-5.6 Sol, while costing 100x less."
The cost gap is concrete. According to the article, "a typical multi-turn search request with gpt-5.6-sol takes >10s and costs ~$0.03 end-to-end." Castform's post-trained 4B model hits the same retrieval accuracy without that per-call bill.
Castform's pipeline runs in three stages on top of Postgres on Neon: corpus storage, synthetic data generation using lakebase_text and lakebase_vector, and RL training plus production inference using Lakebase Search. The company's stated goal is to "make post-training as approachable as prompt engineering."
Why this is a service, not just a model
Castform's own cofounder describes the actual barrier teams face: "Most teams' best training data is just sitting in their databases. The problem is that turning raw data into something usable is hard, and letting agents read, search, and mutate data cheaply at scale requires advanced infra. Pointing Castform at Neon skips both." — Ying Hang Seah, cofounder, Castform.
That sentence names two separate jobs a business can sell: turning a client's raw database into usable training data, and wiring up the infra to train and serve a small model against it. Most teams have the data. Almost none have both the RL expertise and the infra pipeline to turn it into a cheaper working model.
Who buys this
Teams running retrieval, search, or RAG workflows against frontier-model APIs at real volume: support search, internal knowledge lookup, agent tool-calling that hits a search step repeatedly, or any product feature where the same retrieval task runs thousands of times a day against the company's own data.
Bottom line
Castform proved the trade is real: a 4B open model, post-trained on a company's own data, can match a frontier model on retrieval at 100x less cost. The service is turning that proof into a repeatable delivery — take a client's database, generate training data from it, and hand back a cheaper model doing the same job.
Source: https://neon.com/blog/how-castform-neon-beats-frontier-models-on-price-and-efficiency
Tools mentioned
Related Playbooks
Google's TPU 8i Launch Points to a New AI Infrastructure Service: Agent Latency Audits and Inference Rebuilds for Teams Moving Into Multi-Agent Workflows.
Medium · 1-2 weeks to package the first audit offer and land a pilot
The Boring Internal Questions Business Is Still Wide Open. The Real Opportunity Is Private RAG for Teams That Hate Searching.
Medium · 2 weeks to first pilot
Mistral Published 'European AI: a playbook to own it.' The Business Opportunity Is AI Compliance and Procurement Infrastructure for Europe's Single Market.
Medium · 2-4 weeks to first pilot