Manual document processing is now the expensive option.

If your team keys documents into a system by hand — invoices, claims, CVs, product sheets, yours or your clients' — much of that work is now the kind of task a model handles well. How much of yours, and at what accuracy, is what Coldstill measures on your own documents before you change anything. Your team keeps the review step: the people who know the documents check the ones the system flags, rather than keying every field of every document.

The same holds for repetitive work that is not documents at all — transactions categorised by hand, records matched against a list, messages sorted into queues. The calculator below is written in documents because that is the commonest case; the arithmetic is the same for anything decided the same way thousands of times a month.

When this pays for itself

Once keying one document type costs your team more than about $1,000 a month, the arithmetic starts working. At the figures this page uses — 10 minutes a document, $24 an hour — that is roughly 250 documents a month, and the gap widens with every document above it.

It pays hardest on the documents that resist templates: supplier statements, invoices where the line-item detail matters, and the client-specific formats that still end up keyed by hand because nothing off the shelf reads them the way your team does. Those are the ones that stay expensive as volume grows.

The cost comparison

What that looks like on your numbers rather than the example ones.

What manual keying costs, from your numbers

Manual processing, per month $8,000/mo
Fully hosted by Coldstill, from $650/mo

The free read turns this estimate into a written verdict on feasibility and cost, from your own figures, within a working day — send six answers, or see pricing.

Manual figures come from your inputs. In the prefilled example, 2,000 documents a month is about two people's keying at ten minutes per document, and £18 an hour is just under the UK median administrative wage plus statutory employer costs, from official statistics checked 2026-07-31, shown here converted at $1.347 to the pound.

The hosted figure is the full service, not a software licence: Coldstill builds a model fine-tuned to your documents and operates everything — hosting with GPU costs included, accuracy measured against your own historical records before anything switches and re-checked weekly, and the documents the pipeline can't process cleanly flagged for your team to review. At the example volume, the $650/mo minimum works out to about 33 cents a document. If running your volume through a pay-per-use AI service would cost more than twice the minimum, the fee is capped at half that equivalent bill; the exact fee is fixed in the audit.

How switching works

01

Audit

A $650 assessment built from a sample of documents your team has already processed. The verdict covers achievable accuracy, what stays manual, and the running cost. 2–3 days.

02

Build

A model of your own, trained on your historical records where they are consistent enough. Where they aren't, the work runs through a pay-per-use AI service for a short interim period to build a clean set of processed examples, then moves to your own model. The audit sets out which route applies.

03

Run

The model handles the data entry, and your team reviews the documents it flags. Coldstill hosts and operates everything — there is nothing to install and no AI team to hire. Monthly, cancellable, and the trained model is yours if you ever want to take it in-house.

Scope, terms and what's included

If you process documents for clients

Bookkeeping and accounting practices, outsourcers and staffing agencies key documents for many clients at once. That shape is the one the arithmetic works hardest for, and the engagement is built around it.

The free read covers your highest-volume document type across clients, from combined volume figures, within a working day.

Error rates

You will know the model's error rate before anything goes live — field by field, measured against records your own team already processed — and the tool that does the measuring is part of what you receive, so it can be re-run on your side whenever you want. Manual keying has an error rate too. It is simply never measured.

It is not only measured on your documents. The same method is published on two public datasets, with the working shown: 54% fewer mistakes than the strongest frontier model on a routing task, and 1.29× its accuracy on contract fields.

How much review you still need depends entirely on the job, and the two benchmarks are as far apart as they get. On routing, where the answer is one label, 92.2% of messages came back completely correct. On contract extraction, where each document needs ten separate fields right at once, not one document in seventy-six had all ten simultaneously correct — for any model tested, ours included. Both numbers are real and neither is hidden: the more fields a record needs, the more the work is assisted review rather than automation, and the free read tells you which end of that range your task sits at. The full measurements, including where these models still get it wrong.

Your documents are used only to build and evaluate your own model. How accuracy is measured · how your data is handled.

Get a free read on your task

Send what the documents are, roughly how many arrive each month, and how long one takes. The read comes back with feasibility and cost.

The enquiry form takes six questions and a written read comes back within a working day. Email works just as well: hello@coldstill.org.