Choosing between a managed OCR API and self-hosted open source comes down to one comparison: the monthly cost of the API at your volume, against the cost of running the open-source engine yourself, which is mostly the hours your team spends keeping it working. The prices below are public. The hours are yours, and they decide the answer.
This page gives the formula, the published API prices as of October 2026, a worked example and a checklist for the parts a formula does not cover.
What managed OCR APIs publish
All prices are US dollars at list price, per 1,000 pages, as read on 1 October 2026. Check the linked page before you decide, since prices change.
| Service | Price per 1,000 pages | Free tier |
|---|---|---|
| Textract, text | $1.50 ($0.60 above 1M) | 1,000 pages/month for 3 months |
| Textract, tables | $15 ($10 above 1M) | See pricing page |
| Azure Read | $1.50 ($0.60 above 1M) | 500 pages/month |
| Google OCR | $1.50 ($0.60 above 5M) | First 1,000 pages |
| Mistral OCR 4.1 | $4 | See pricing page |
Azure's pricing page did not display figures when we read it; the Azure figures come from Microsoft's Retail Prices API for the East US region. Mistral's batch discount lowers its price further, and many older articles still quote the price of its previous model, so check the current page. Textract and Azure also charge more for tables, forms and prebuilt extraction than for plain text, so compare the same capability across services.
What self-hosting costs
Two parts of the cost have a price you can look up, and the rest does not.
Infrastructure is the part with a public price. As an example, an AWS c6i.xlarge instance (4 vCPUs, 8 GiB) costs about $0.17 an hour on demand in us-east-1, close to $124.80 a month, according to the Vantage instance table, which mirrors AWS list prices; confirm it on the AWS pricing page. Block storage on gp3 is $0.08 per GB-month on the EBS pricing page. A GPU instance such as the g5.xlarge (one NVIDIA A10G) is about $1.006 an hour, close to $734 a month, per the same table.
The rest has no public price: the engineering time to build it, preprocessing and post-processing code, upgrades, monitoring and on-call, labelling a test set, security work, and a fallback for pages that fail. These are real costs, and they are the reason the decision is not just a comparison of price lists.
The formula
With P as pages per month and p as the API price per 1,000 pages, the API costs P / 1000 × p a month. Self-hosting costs the infrastructure plus the people: infrastructure and storage, plus the monthly hours of upkeep times your loaded hourly rate, plus the build hours and the labelling, spread over the months you expect to run it.
def monthly_cost_api(pages: int, price_per_1000: float) -> float:
return pages / 1000 * price_per_1000
def monthly_cost_self_hosted(infra: float, upkeep_hours: float, hourly_rate: float,
build_hours: float = 0, labelling_cost: float = 0, months: int = 12) -> float:
"""Infrastructure, upkeep, and the one-off build and labelling spread over `months`."""
one_off = build_hours * hourly_rate + labelling_cost
return infra + upkeep_hours * hourly_rate + one_off / months
def break_even_rate(pages: int, price_per_1000: float, infra: float, upkeep_hours: float,
build_hours: float = 0, labelling_cost: float = 0, months: int = 12):
"""Hourly rate at which both options cost the same, or None when there is no labour to price."""
labour_hours = upkeep_hours + build_hours / months
if labour_hours == 0:
return None # compare infrastructure and labelling with the API cost directly
return (monthly_cost_api(pages, price_per_1000) - infra - labelling_cost / months) / labour_hoursThe loaded hourly rate is your number: salary, overhead and the opportunity cost of what else those hours could do. This page does not guess it. The code sets the build hours and the labelling cost to zero by default, so a call that leaves them out covers ongoing costs only and understates self-hosting; pass your own values to include the one-off work. With no upkeep and no build hours there is no labour to price, so break_even_rate returns None, and the comparison is just infrastructure against the API. A negative result means self-hosting costs more than the API even if the labour were free.
A worked example
Take 100,000 pages a month of plain text OCR. At $1.50 per 1,000 pages the API costs $150 a month, and at Mistral OCR 4.1's $4 it costs $400. Self-hosting on the c6i.xlarge with 100 GB of storage costs about $132.80 a month in infrastructure ($124.80 plus $8.00).
The infrastructure alone is almost the same as the cheapest API. So the question is how many hours it takes to keep the engine running. As an illustrative placeholder, assume 20 hours a month of upkeep. Then self-hosting is cheaper than the $1.50 API only if one hour of that work costs less than about $0.86 (the $17.20 difference divided by 20 hours). At that price of labour, self-hosting does not win. Against the $4 API the same arithmetic gives about $13.36 an hour, which is where your own rate starts to matter.
infra = 124.80 + 100 * 0.08 # c6i.xlarge plus 100 GB of gp3 storage
print(monthly_cost_api(100_000, 1.50)) # 150.0
print(round(infra, 2)) # 132.8
print(round(break_even_rate(100_000, 1.50, infra, 20), 2)) # 0.86A GPU changes the picture only at high volume. A g5.xlarge at about $734 a month costs less than a $1.50 API only above roughly 489,000 pages a month, and less than a $4 API above roughly 183,500 pages a month, before any labour. These numbers are illustrations with sourced prices and placeholder hours; replace the hours and the rate with your own.
What the formula leaves out
- Document types. Tesseract documents its own limits on skew, low resolution and tables, and a managed service may handle them differently. Measure on your files.
- Data residency. Self-hosting keeps data in your network. Managed APIs often offer region choices, so check each vendor.
- Latency and throughput. Measure them, and note that some vendors price batch and synchronous calls separately.
- Team capacity. Someone has to own preprocessing, upgrades and monitoring.
- Accuracy target. Test 50 of your own documents on each option: pick a mix that covers your best, median and worst scans, label the fields you care about, run all candidates on the same inputs and set the pass threshold before you look at results. The scoring code is in Is Tesseract accurate enough for production?.
Where anyformat fits
anyformat is a third option: a platform that returns typed fields with a confidence score and the source text and page, rather than raw OCR text. Its self-hosted and on-premise deployment is an Enterprise option with no public price, so it has no row in the formula above; ask for a quote if that is your case, and run the same 50 documents through it. See on-premise and air-gapped document extraction for how that deployment works.
Frequently asked questions
Is it cheaper to self-host OCR?
Only at high volume or when the hours to run it are small. At $1.50 per 1,000 pages the infrastructure costs about as much as the API for 100,000 pages a month, so labour decides.
How much does a managed OCR API cost?
Published list prices for plain text OCR were $1.50 per 1,000 pages at AWS Textract, Azure Document Intelligence and Google Document AI, and $4 at Mistral OCR 4.1, as of 1 October 2026. Tables, forms and prebuilt extraction cost more.
When does self-hosting make sense beyond cost?
When data cannot leave your network, when you need to run offline, or when you need control over the model and its version.
What is the break-even volume?
It depends on your hours and hourly rate. Use the formula on this page: the loaded hourly rate at which both options cost the same is the API cost minus the infrastructure, divided by the monthly hours.
How do I compare accuracy between options?
Test 50 of your own documents, score the fields you care about, run every candidate on the same inputs and set the threshold first.







