A document parsing API turns a PDF, scan or image into structured text you can use in code: markdown or JSON with headings, tables and reading order preserved, and sometimes the position of every block on the page. It is the first step before extraction, retrieval or an agent reads the document, and the differences between providers sit in price, where the data can run, and what comes back besides the text.
This page compares seven APIs on four things you can check on each vendor's public pages: the published starting price and free tier, where it can be deployed, the compliance it lists, and what the output carries beyond markdown. Prices are what each vendor published on 30 September 2026, in the currency it publishes, and we did not convert them. The tools are listed in alphabetical order. We do not rank them, because no benchmark covers all seven and the ones that exist are run by a vendor with a stake in the result.
anyformat is one of the seven. We built it, so read its section with that in mind; it is described with the same criteria and the same limits as the others.
| Tool | Published price for parsing | Free tier | Deployment | Output beyond markdown |
|---|---|---|---|---|
| anyformat | Parse: 25 credits per page, 10 credits = €0.01 (€0.025 per page, €25 per 1,000 pages) | 50,000 credits on signup, one time | EU-native cloud, private cloud, on-premise including air-gapped | Block layout; per-field confidence and evidence on extraction |
| Extend | 2 credits per page on the performance tier, $0.0125 per credit pay-as-you-go | Free credit allowance | US1, US2 and EU1 regions, VPC and hybrid | Review queue and evals |
| LandingAI ADE | About 1.5 credits per page, 100 credits = $1 | 1,000 credits | US and EU hosting, VPC and on-premise on Enterprise | Bounding boxes through its Ground API |
| LlamaParse | 1 to 45 credits per page by tier, 1,000 credits = $1.25 | 10,000 credits | US cloud and an EU region, enterprise VPC | Layout and confidence on some tiers |
| Mistral OCR | Per 1,000 pages; the figure was not shown on the pages we checked | Not shown on the pages we checked | EU company, API | Paragraph-level boxes, block labels and block confidence (OCR 4.1) |
| Reducto | $10 per 1,000 pages on its r-1 model | $150 in credits | EU and AU regions on Growth, VPC and on-premise on Enterprise | Layout and bounding boxes |
| Unstructured | $15 per 1,000 pages after 10,000 free pages | 10,000 pages | SaaS, dedicated instance, in-VPC, bare metal on Business | Element arrays and chunks |
Sources for the table, all read on 30 September 2026: Extend, LandingAI, LlamaParse, Mistral (pricing and models), Reducto, Unstructured and anyformat's own pricing page. Where a figure was not on those pages, the table says so instead of guessing.
1. anyformat
anyformat is a document extraction platform built around a schema: you define the fields, and the platform parses, extracts, classifies, splits and validates, with a review step for what falls below a confidence threshold. Parsing is available on its own through the v3 API, which returns markdown, and through an MCP server for agents.
Parse costs 25 credits per page, and 10 credits are €0.01, so €0.025 per page or €25 per 1,000 pages; some anyformat pages quote that same figure as $25 per 1,000 pages, but the credit price is set in euros. Extract is 35 credits per page. The free tier grants 50,000 credits on signup, one time. anyformat is EU-native and deploys as cloud, private cloud or on-premise, including air-gapped, which is in production at a government customer; there is more in on-premise and air-gapped document extraction. It lists ISO 27001 and GDPR, and a zero-retention mode is available on request.
What sets it apart is what happens after parsing: extracted fields come back with a calibrated confidence score and the source text and page they were read from, which feeds the review loop. The trade-offs are real. Parsing at €25 per 1,000 pages costs more than Reducto's r-1 model, and the Edit operator is still in beta. Benchmark detail is in the anyformat benchmark post, which we published ourselves, and the pricing is on the pricing page. For a direct comparison with each tool below, see the /vs pages linked in the sections that follow.
2. Extend
Extend offers parse, extract, split, classify and edit APIs with a visual Studio, and it adds human-in-the-loop review and evaluation tooling.
It prices in credits: pay-as-you-go is $0.0125 per credit, with a Scale plan at $500 per month that includes 50,000 credits at $0.01, and the performance parsing tier uses 2 credits per page. Regions are US1, US2 and EU1, with bring-your-own-cloud, VPC and hybrid options, and we found no fully air-gapped option in its documentation. It lists SOC 2 Type II, HIPAA (a business associate agreement on the Scale plan) and GDPR. It is a US company. Some options multiply the credit cost, for example priority processing, so check the multiplier for the tier you would use. See anyformat vs Extend. Source: Extend's pricing page, read 30 September 2026.
3. LandingAI ADE
LandingAI's Agentic Document Extraction is a vision-first parser that returns markdown and structured fields, with split and extract operations on top.
A dollar buys 100 credits, a typical page is about 1.5 credits, and the free allowance is 1,000 credits. The Team plan starts at $250 per month. Hosting in the EU is listed on every tier at the same price as the US, VPC and on-premise are Enterprise options, and it lists SOC 2 Type II and HIPAA for Team and above. Its Ground API returns bounding boxes for extracted content. Two caveats from its public documentation: the default DPT-3 Pro parser does not return a confidence score, and with zero data retention enabled the Ground API is not available. See anyformat vs LandingAI. Source: LandingAI's pricing page, read 30 September 2026.
4. LlamaParse
LlamaParse, from LlamaIndex, covers parse, extract, classify and split on one credit wallet,. LlamaIndex also publishes LiteParse, an open-source local parser that outputs plain text.
A thousand credits cost $1.25, and a page costs from 1 credit on the Fast tier to 45 on Agentic Plus, so $1.25 to $56.25 per 1,000 pages, with extra credits for pages containing forms. The free tier is 10,000 credits. It runs in the US and in an EU region live since July 2026, with VPC for enterprise. LlamaParse entries lead ParseBench, a parsing benchmark that LlamaIndex itself runs (leaderboard of 21 September 2026), so read that ranking knowing who publishes it. Its focus is parsing rather than review workflows. See anyformat vs LlamaParse and LlamaParse alternatives. Sources: LlamaIndex's pricing page and documentation, read 30 September 2026.
5. Mistral OCR
Mistral OCR is a single OCR model behind an API rather than a platform: you send a document and get markdown back. The current model is OCR 4.1, which returns paragraph-level bounding boxes, structural block labels and block-level confidence scores; OCR 4.0 was deprecated on 29 September 2026 and retires on 30 September 2026, so use 4.1.
Mistral is a French company, and its API is priced per 1,000 pages on its pricing page; the docs and pricing pages did not show the figure when we checked, so we do not quote one, and we found no compliance certifications listed on the pages we checked. The scope is narrower than the others here: there is no workflow builder, field extraction or review step, and you build those around it. For Spanish-language documents, see the best OCR APIs for Spanish documents. Sources: Mistral's models page and pricing page, read 30 September 2026.
6. Reducto
Reducto is a code-first parse and extract API with a Studio for testing. Its r-1 model, launched in September 2026 and still a preview, is priced at $10 per 1,000 pages for parsing and $20 for extraction on its public pricing page.
Free credits start you at $150. It offers EU and AU regions on the Growth plan and VPC or on-premise deployment for Enterprise, and it lists SOC 2, HIPAA and zero data retention. It also publishes RD-TableBench, an open table-extraction benchmark. The r-1 model is opt-in and in preview, so check the billing for the model you call. See anyformat vs Reducto. Source: Reducto's pricing page, read 30 September 2026.
7. Unstructured
Unstructured is an ingestion platform with a partly open-source core: it turns documents into typed elements and chunks ready for a vector store, with an extract step that takes a JSON schema.
After 10,000 free pages it charges $0.015 per page, which is $15 per 1,000, with custom pricing on Business. Deployment options include SaaS, a dedicated instance, in-VPC and bare metal on Business. It lists HIPAA, SOC 2 Type 2, GDPR and ISO 27001, and FedRAMP High as available. Its output is element arrays rather than per-field values with confidence, which suits ingestion pipelines more than review workflows. See anyformat vs Unstructured. Source: Unstructured's pricing page, read 30 September 2026.
How to choose
- You need on-premise or VPC deployment: anyformat, or Reducto, LandingAI or Unstructured on their higher tiers.
- You need air-gapped specifically: anyformat states it, in production at a government customer. We did not find it stated for the others on the pages we checked, so ask each vendor directly.
- You need EU data residency: anyformat is EU-native, and Extend, LandingAI, LlamaParse and Reducto offer an EU region.
- You feed a RAG index and want the lowest friction: LlamaParse or Unstructured.
- You want a model, not a platform, and will build the rest: Mistral OCR or Reducto.
- You need review queues and confidence on fields, not only text: anyformat or Extend.
- You want bounding boxes from a single endpoint: LandingAI's Ground API or Mistral OCR 4.1.
Whatever you pick, run it on your own hardest documents before you commit. A sample of twenty real files will tell you more than any roundup, including this one, and it shows how each option handles the tables and scans that break parsers. The same test works for every tool above.
Frequently asked questions
What is a document parsing API?
An API that converts a document into structured text: markdown or JSON with headings, tables and reading order preserved, and often the position of each block. Extraction, which pulls out named fields, usually runs on top of parsing.
What is the difference between parsing and extraction?
Parsing reads the document and returns its content and structure. Extraction takes that content and returns the specific fields you defined in a schema, such as an invoice number or a total.
Which document parsing API is the most accurate?
No benchmark covers all seven tools, and the main public ones are run by a vendor. We do not rank them for that reason. Test on your own documents.
Which document parsing APIs can run on-premise?
Among these seven, anyformat offers on-premise including air-gapped. Reducto and LandingAI offer on-premise on Enterprise plans and Unstructured offers in-VPC and bare metal on Business, but we did not find air-gapped operation stated for them. Extend offers VPC and hybrid, and we found no air-gapped option in its documentation. Confirm with each vendor, since these options change.
Is there a free document parsing API?
Most of these have a free allowance. anyformat grants 50,000 credits on signup, LlamaParse 10,000 credits, Unstructured 10,000 pages, LandingAI 1,000 credits and Reducto $150 in credits. Mistral's free terms were not shown on the pages we checked.

