For software vendors and SaaS selling to enterprise and public sector

Ship the AI feature.Run it inside thecustomer's deployment.

Your enterprise and government customers won't accept OpenAI in the data flow, and per-token costs don't fit a fixed-price subscription. We build a compact model for the one feature you need, license it for redistribution, and hand you the pipeline. It runs on the customer's server, offline.

0
third-party AI in the data flow
1 license
redistributable, permissive base
2-4 wks
from spec to a measured pilot
Yours
weights, pipeline, eval harness

Why a company wants its own model

What it gets done for the business, not just for IT.

01

Win and keep the deals that forbid third-party AI

Answer the security questionnaire line about AI subprocessors with "none". The model ships inside your product.

02

Add AI without adding a per-token COGS line

A flat fee per task, once. Your margin on the subscription stays yours as usage grows.

03

Keep the roadmap independent of any AI vendor

No deprecations, no price changes, no terms-of-service surprises. Retrain on a better open base whenever you choose, with the recipe we hand over.

What your model will do instead of people

Each line is a candidate for a first pilot: one task, one metric, three weeks. Each task gets its own compact model.

All use cases →

What gets in the way today

And what we do about it.

Enterprise security reviews reject OpenAI, Anthropic and Google in the data flow
The model runs inside your product, inside the customer's deployment. There is no external call to review.
Per-token costs are unpredictable inside a fixed-price subscription
No tokens. One flat fee for the build; inference costs the customer's electricity.
We have no ML engineers and no time to become ones
Two of ours build the feature end to end. You get a container, a model, an eval harness and a retraining recipe your backend team can run.
Can we even redistribute a model to customers?
Yes. We build on permissively licensed open-weights bases (Apache-2.0 / MIT) and deliver a license summary with every artifact.
Our customers run on CPU-only servers
Most first features (classification, extraction, routing) run on CPU. We size and benchmark on your reference environment.

What is different for software vendors

Built to ship inside your product, not just inside one company.

How your model comes to be

We start from your task and your success metric, build the data, train your own compact model, prove it on a gold set, and install it on your server. The model, the data and the code stay with you.

  1. Step 1
    Define the task

    One or two narrow tasks with a clear success metric. We write the eval first.

  2. Step 2
    Build the data

    Our engineers collect, clean and label thousands of examples from your real inputs, using the strongest available AI tooling, or fully offline if required.

  3. Step 3
    Train your model

    We take a compact open-weights model and train it for your task, and only your task, on your data. Measured against the gold set until it clears the bar.

  4. Step 4
    Install on your server

    Delivered as a container with an OpenAI-compatible API, on your server, offline. We hand over weights, data and pipeline: the model is yours.

Our approach →

Purpose-built beats general-purpose on narrow tasks

Public results, not ours, from teams that replaced a frontier API with a small model built for one job.

97% vs 88%

Checkr: fine-tuned Llama-3-8B vs GPT-4 on background-check classification. ~$800/mo instead of $7-12K, 0.5 s instead of 15 s.

25 of 27

Predibase "LoRA Land": fine-tuned 7B adapters matched or beat GPT-4 on 25 of 27 tasks, each trained for under $8 of GPU time.

Gartner predicts that by 2027 organizations will use small, task-specific models three times more than general-purpose LLMs.

Sources and details on the FAQ page. Your numbers will differ, that's what the pilot measures.

Your own local model vs. a cloud API

Frontier cloud APIYour model on your server
Where data goesVendor's servers, another jurisdictionStays inside your network
Cost modelPer token, forever, vendor-pricedOne flat fee + a server you already own
LatencySeconds, internet-dependentSub-second, on-prem
Accuracy on your taskGood, genericUsually equal or better, it was built for it
Who owns itThe vendor; you rent accessYou: weights, data, code
Vendor riskModel deprecations, price changes, ToS changesNothing to deprecate, no external calls
Works offlineNoYes
Open-ended chat, coding, reasoningExcellentNot the goal, one job, done well

One flat fee per task

From $7,000 to train and deploy your model. Two engineers, data, training, verification, installation on your server, and a follow-up fine-tune one month after go-live. No per-token pricing, no subscription: the model is yours, using it costs nothing extra.

For vendors, the flat fee covers the feature once, for all your customers. Multi-tenant variants, several languages and CI integration move the number up; see the pricing page.

Pilot & pricing →

We also train models for

Same method, different jobs. Pick the page that speaks your language.

For
Regulated mid-market

Finance, insurance, legal and professional services: your own AI model for the work you can't send to a cloud API.

For
Healthcare & health-tech

Clinics, payers, RCM and health-tech: clinical and claims paperwork without PHI ever leaving your environment.

Run a pilot on your data

Two to four weeks. One task. A measurable result before you commit.

Start a pilot →