Document parsing for AI agents means giving an agent a tool that turns a PDF, scan or image into text it can read, with the structure kept: headings, tables and reading order. An agent cannot look at a file the way a person does. It reads what a tool returns, so whether it can work with your documents depends on what that tool sends back.
MCP, the Model Context Protocol, is the standard way for an agent to call such tools. A server advertises tools, the agent chooses when to call them, and the results come back as content the model reads. This page explains what that means for documents, describes how providers expose it, and walks through connecting a coding agent to a parsing server and parsing a PDF.
Why an agent needs a parsing tool
A tool call carries JSON, not files. According to the MCP specification, tool results are text, images, audio, resource links or embedded resources, and the model reads them as content. A PDF therefore has to reach the server some other way, through a URL or an upload, and come back as text. anyformat's own MCP guide describes the same constraint: the agent stages the file or passes a URL, and the parse result returns as markdown.
Two costs follow from that. The first is payload size. In anyformat's documentation example for a real invoice run, the parse section is about 90 percent of the payload, which is why the API lets the caller ask for no parse output when only the fields matter; that is one example, not a measured average. The second is the tool list itself. Every tool an agent sees takes space in its context, so some providers let you filter tools or split them into smaller servers.
How other providers expose agent access
Providers follow two patterns, as of 30 September 2026. A hosted remote server is added to the agent by URL and authenticated either with OAuth, where the client opens a browser to sign you in, or with an API key sent as a header; Reducto, Extend and LlamaIndex publish servers like this. A local server runs on your machine and is started by the client over standard input and output; depending on the provider it needs an API key or none, as with the open-source Docling server. Not every provider has one: for Mistral we found an API and an agents framework that can call MCP tools, but no first-party server for parsing documents.
Setting up any of them follows the same steps. Get credentials if the server needs them, add the server to your client using the provider's own instructions, ask the client to list the tools and try a small synthetic document. What differs is what matters when you choose: OAuth or API key, hosted or local, and how many tools the server puts in front of the agent. Check each provider's documentation for current details, because these servers change quickly.
Connect an agent to anyformat
anyformat exposes one remote MCP server at https://api.anyformat.ai/mcp. It uses the Streamable HTTP transport and authenticates with your API key as a Bearer token, the same key as the REST API. There is no OAuth flow today, so the custom connector panel in Claude's web and desktop apps, which expects OAuth, is not the way in. Use a client that can send a header, such as Claude Code, Cursor, Codex or OpenCode, or an mcp-remote bridge. The documentation describes the server in one line: anyformat speaks MCP, and you connect your agent in one config block.
Create a key at app.anyformat.ai, export it, and add the server. For Claude Code, from the documentation:
export ANYFORMAT_API_KEY="af_..."
claude mcp add --transport http anyformat https://api.anyformat.ai/mcp \
--header "Authorization: Bearer $ANYFORMAT_API_KEY" --scope user
claude mcp listFor Cursor, add this to ~/.cursor/mcp.json or .cursor/mcp.json:
{
"mcpServers": {
"anyformat": {
"url": "https://api.anyformat.ai/mcp",
"headers": { "Authorization": "Bearer ${env:ANYFORMAT_API_KEY}" }
}
}
}The documentation has the equivalent blocks for Codex and OpenCode. Keys carry scopes: read tools need read access, tools that create, run, update or delete need write access, and a tool the key cannot call does not appear in the list. Delete tools are marked destructive and ask the agent to request approval first.
Parse a PDF with an agent
Use a synthetic PDF for this, not a client document: the agent saves the output to disk and the call is billed.
With the server connected, you ask in plain language:
Parse invoice.pdf with anyformat and save the markdown to invoice.md.Behind that request the agent follows the quick-parse path in the documentation. It stages the file, which returns an upload slot with a short object id, and sends the bytes to that slot with a shell command, since no MCP tool carries file bytes. If the PDF is already at a public HTTPS URL, it skips the staging and passes the URL. It then calls parse_document with the object id, a fast parse of one document, and get_run to wait for the result. The markdown comes back in the run's parse output, and the documentation says that output is not retained, so the agent has to write it to a file, which is why the prompt says to save it. The documentation's sample invoice takes about 13 seconds.
From there the markdown is an ordinary file the agent can read, search and quote. If you need fields instead of text, the same server can create a workflow with a parse node and an extract node and run documents through it, which returns each field with a confidence score and the source text and page. That is covered in the MCP guide, along with the Knowledge path for asking questions across documents, which is in alpha.
What to watch
- Credits. Parsing through a workflow or
parse_documentis a billed call. Check the credit table on the pricing page and use a test key with a limit for experiments. - Write access. A key with write scope can create, run and delete. Give agents the smallest scope that does the job.
- Real documents. Anything the agent parses ends up in its context and possibly in its logs. Keep client documents out of early tests.
- Volume. MCP suits an agent working through a handful of documents. A batch of thousands belongs in the REST API or an SDK, called from code.
Frequently asked questions
What is an MCP server for document parsing?
A server that advertises tools an AI agent can call to convert documents into text or structured data. The agent decides when to call them, and the results come back as content the model reads.
Can an AI agent read a PDF directly?
Not through a tool call alone, because a call carries JSON and not file bytes. The PDF has to reach a server as a URL or an upload, and the parsed text comes back as the result.
Do I need OAuth to connect an agent to anyformat?
No. anyformat uses an API key as a Bearer header and has no OAuth flow today. That means you connect it from clients that can set a header, or through a bridge, rather than from a connector panel that requires OAuth.
How do I stop an agent from deleting workflows?
Use a key with read scope only, or keep write scope for the agent that needs it. Delete tools are marked destructive and ask the agent to request approval first.
Is MCP better than the REST API for parsing?
For an agent deciding what to do next, MCP. For fixed pipelines and large batches, the REST API or an SDK called from your own code.







