Stop paying frontier prices for a task that doesn't need a frontier model.

One repetitive judgement your business makes thousands of times a day — fields read off documents, tickets routed to the right queue, records classified or matched — moved onto a small model trained on the answers your team has already given. It is measured against your current system before anything switches, and the model is yours outright.

The qualifying test is simple: you already have a history of correct answers. That history is the training data, so there is nothing to annotate and nothing to specify from scratch. Six questions gets you a written read, free, within a working day.

It is measured, not asserted

Two public tasks, same frontier model on the other side of both, scored the same conservative way. Neither number is a projection.

Reading fields off contracts 1.34× the accuracy of Claude Opus 5, given the same worked examples
Routing messages to 77 queues 54% fewer mistakes than the same frontier model
Cost per document 22×–62× less at the providers' standard rates
Malformed records 0 across both runs, against 13–79 per 1,000

The full measurements — every field, both baselines, the method, and the places these models still get it wrong.

How it works

01

Audit

An assessment of one endpoint: whether it can move to a more efficient model, the projected saving, and the risks. 2–3 days, and it behaves as a deposit rather than a fee.

02

Migrate

A model trained on your production logs, measured against your current one on held-out data. Nothing switches until it matches, and the harness that proves it ships to you.

03

Run

Coldstill hosts and operates it — GPU, monitoring, weekly quality checks and retraining. You point your calls at a new endpoint and nothing else on your side changes.

What it costs

StageFeeWhat it buys
Audit $800 A written verdict in 2–3 days. The whole fee comes off the migration if you go ahead within 90 days, and comes back in full if the task turns out to be one Coldstill can't take on.
Migration $4,000 Fixed fee, 2–3 weeks. Trained weights, serving config, the eval harness and the handover doc — all owned by you.
Hosted, optional from $650/mo Capped at half your current or equivalent API bill. Monthly, and cancelling it does not cost you the model.

No hourly billing and no infrastructure to buy. Full scope and terms, or start with the free read.

Get a free read on your task

Answer these and a written read comes back within a working day: your current spend with the arithmetic shown, whether the task can move to a model you own, the likely saving, and the risk most likely to threaten it. If the honest answer is that your current setup already does the job, that is what it says.

How is the work done today?

Numbers, not documents — nothing confidential is needed, and what you send goes to one mailbox and no further. What happens to it.

Prefer email? [email protected] — the same answers in a message work just as well.