For software vendors and SaaS selling to enterprise and public sector
Your enterprise and government customers won't accept OpenAI in the data flow, and per-token costs don't fit a fixed-price subscription. We build a compact model for the one feature you need, license it for redistribution, and hand you the pipeline. It runs on the customer's server, offline.
What it gets done for the business, not just for IT.
Answer the security questionnaire line about AI subprocessors with "none". The model ships inside your product.
A flat fee per task, once. Your margin on the subscription stays yours as usage grows.
No deprecations, no price changes, no terms-of-service surprises. Retrain on a better open base whenever you choose, with the recipe we hand over.
Each line is a candidate for a first pilot: one task, one metric, three weeks. Each task gets its own compact model.
Tickets, documents, events, into the customer's own taxonomy.
Strict JSON matching your schema, with confidence per field.
Fully offline, with citations, in the customer's language.
For DLP, compliance and redaction features.
In your product's DSL or your customers' SQL dialect.
Into fixed templates your UI already renders.
And what we do about it.
Built to ship inside your product, not just inside one company.
We start from your task and your success metric, build the data, train your own compact model, prove it on a gold set, and install it on your server. The model, the data and the code stay with you.
One or two narrow tasks with a clear success metric. We write the eval first.
Our engineers collect, clean and label thousands of examples from your real inputs, using the strongest available AI tooling, or fully offline if required.
We take a compact open-weights model and train it for your task, and only your task, on your data. Measured against the gold set until it clears the bar.
Delivered as a container with an OpenAI-compatible API, on your server, offline. We hand over weights, data and pipeline: the model is yours.
Public results, not ours, from teams that replaced a frontier API with a small model built for one job.
Checkr: fine-tuned Llama-3-8B vs GPT-4 on background-check classification. ~$800/mo instead of $7-12K, 0.5 s instead of 15 s.
Predibase "LoRA Land": fine-tuned 7B adapters matched or beat GPT-4 on 25 of 27 tasks, each trained for under $8 of GPU time.
Gartner predicts that by 2027 organizations will use small, task-specific models three times more than general-purpose LLMs.
Sources and details on the FAQ page. Your numbers will differ, that's what the pilot measures.
| Frontier cloud API | Your model on your server | |
|---|---|---|
| Where data goes | Vendor's servers, another jurisdiction | Stays inside your network |
| Cost model | Per token, forever, vendor-priced | One flat fee + a server you already own |
| Latency | Seconds, internet-dependent | Sub-second, on-prem |
| Accuracy on your task | Good, generic | Usually equal or better, it was built for it |
| Who owns it | The vendor; you rent access | You: weights, data, code |
| Vendor risk | Model deprecations, price changes, ToS changes | Nothing to deprecate, no external calls |
| Works offline | No | Yes |
| Open-ended chat, coding, reasoning | Excellent | Not the goal, one job, done well |
From $7,000 to train and deploy your model. Two engineers, data, training, verification, installation on your server, and a follow-up fine-tune one month after go-live. No per-token pricing, no subscription: the model is yours, using it costs nothing extra.
For vendors, the flat fee covers the feature once, for all your customers. Multi-tenant variants, several languages and CI integration move the number up; see the pricing page.
Same method, different jobs. Pick the page that speaks your language.
Two to four weeks. One task. A measurable result before you commit.