Converting a PDF to markdown for an LLM means turning pages into text that keeps the structure a model needs: headings as headings, tables as tables, columns in reading order. Markdown is the usual target because models read it well, it is usually lighter than HTML, and a RAG pipeline can split it on headings instead of on arbitrary character counts.
You do this for one of three reasons: to index documents for retrieval, to paste a document into a prompt without losing its tables, or to give an agent a file it can read. Plain text extraction is not enough for any of the three, because a table flattened into one cell per line is worse than no table.
This page runs three open-source converters on the same PDFs, shows the code, and reports what came out. It also says what we did not test, which includes anyformat.
What you will need
- Python 3.10 or newer and about 4 GB of free disk if you install all three, because Docling and Marker download models on first use.
- One PDF of your own, ideally a sample or fictional document rather than a real client file. Use the ugliest one you have: a table, two columns, or a scan. Converters look the same on clean single-column text.
- A terminal. The three open-source tools need no API keys.
- One virtual environment per tool. We installed each one separately, so their dependencies do not collide:
pip install "markitdown[pdf]" # environment 1
pip install docling # environment 2
pip install marker-pdf # environment 3How we tested
We generated three fictional PDFs: a two-page report (headings, a bulleted list, a 5 by 4 table with a header row, then a page in two columns), a scan-style PDF with no text layer (a slightly rotated image of three sentences and a small table), and a 40-page document. We ran the minimal snippet from each project's README on an Apple Silicon Mac with no GPU, on 2026-09-30. The versions were MarkItDown 0.1.8, Docling 2.131.0 and Marker 2.0.0.
This is one machine and small synthetic files, each timed once. Treat it as a way to see the failure modes, not as a benchmark.
Step 1. Run the three converters
Each one is a few lines. MarkItDown, from Microsoft:
from markitdown import MarkItDown
text = MarkItDown().convert("report.pdf").text_contentDocling, from IBM Research:
from docling.document_converter import DocumentConverter
text = DocumentConverter().convert("report.pdf").document.export_to_markdown()Marker, from Datalab. Put it under a main guard, because on multi-page PDFs it starts worker processes and, without the guard, re-runs the top of your script:
from marker.converters.pdf import PdfConverter
from marker.models import create_model_dict
from marker.output import text_from_rendered
if __name__ == "__main__":
converter = PdfConverter(artifact_dict=create_model_dict())
text, _, images = text_from_rendered(converter("report.pdf"))What we saw
| MarkItDown | Docling | Marker | |
|---|---|---|---|
| Table | Broken: one cell per line, columns mixed | Correct | Correct |
| Two columns | Correct order, lines hard-wrapped | Correct order, paragraphs unwrapped | Correct order, one heading dropped |
| Heading levels | None kept | All headings came out as ## |
## or ###, levels flattened |
| Scanned page | Empty output | Read, table correct | Crashed without an extra install (below) |
| 40 pages, second run | 0.5 s | about 3 s | about 3.5 s |
| First run | none | 41 s, plus 28 s to initialise | 36 s, then about 7 s |
| Install and models | 150 MB | 1.1 GB plus about 570 MB of models | about 1 GB plus 418 MB of models |
Times come from one Apple Silicon Mac with no GPU, single runs, on 2026-09-30.
MarkItDown is fast and only good on clean text
It converted two pages in 0.04 seconds and kept the two-column reading order. It kept no headings, turned bullet glyphs into (cid:127) lines, and scrambled the table: each cell became its own paragraph and the column order was mixed.
Region Q1 Q2 Q3
North
100
120
140
South
On the scanned PDF it returned an empty string. MarkItDown's README describes a separate markitdown-ocr plugin that sends images to an LLM with vision; we did not test it. For clean single-column text it is the cheapest option by a wide margin. For tables or scans it is not enough on its own.
Docling gave the best output here, at a cost
Docling returned the table as a correct markdown table, the list as a list and the two columns in order with paragraphs unwrapped. It also ran OCR on the scan with no configuration, returning the three sentences and a correct table. We compared by eye against the source and saw no errors; that is an observation, not an accuracy score.
| Region | Q1 | Q2 | Q3 |
|----------|------|------|------|
| North | 100 | 120 | 140 |
| South | 107 | 125 | 143 |
The cost is about 1.1 GB of packages and about 570 MB of models, roughly 70 seconds the first time (model download and initialisation), then a few seconds per document. It also flattened heading levels: every heading came out as ##, so a title and its subsections look the same. If your chunker splits on heading depth, that matters.
Marker: check your install before you trust it
Marker read the text PDFs, with a correct table and correct two-column order, apart from one heading it dropped on the two-column page. Two things went wrong in our run, and both are worth knowing.
First, on a Mac without a GPU, Marker 2.0.0 could not convert the scanned page. It stopped with SpawnError: llama-server binary not found. Its README says the CPU and Apple Silicon mode needs the llama-server binary from llama.cpp installed separately. We did not install it, so we have no Marker result on scans and we are not claiming one.
Second, on a 40-page PDF where every section had identical body text, Marker returned 701 characters and dropped the body of sections 2 to 40. We suspect its repeated-header removal, but we did not confirm that, so treat it as something we observed, not explained. On a 40-page PDF with unique text per section all 40 sections survived, but only 42 of 81 headings came through, with levels flattened to ### and ####.
Read the licenses before you ship any of them. Marker's README says the code is Apache 2.0 and that the model weights use "a modified AI Pubs Open Rail-M license (free for research, personal use, and startups under $5M funding/revenue)", with a commercial route beyond that through the vendor. Docling's code is MIT, and its README sends you to each model's own license. MarkItDown is MIT.
What none of them gave us
No converter here preserved true heading levels, and none returns where on the page a block came from or how sure it is of it. For a RAG chunker that is usually fine. For anything where a reviewer must see the source, or where you decide automatically what to trust, you need that information, which is what the anyformat section below describes.
Check on your own PDFs
Do not pick from our table. Write down a few things you know are in your document and test for them. It takes ten minutes and tells you more than any benchmark. The function below reports only which check failed, never the document's content, so its output is safe to log.
def check_markdown(md: str, must_contain: list[str], table_row: list[str]) -> list[str]:
problems = []
for i, needle in enumerate(must_contain, start=1):
if needle not in md:
problems.append(f"expected text #{i} is missing")
# a table row survives if a markdown line starting with | has exactly these cells
def is_table_row(line: str) -> bool:
cells = [c.strip() for c in line.strip().strip("|").split("|")]
return line.lstrip().startswith("|") and cells == table_row
if not any(is_table_row(line) for line in md.splitlines()):
problems.append("table row is not intact")
return problems
print(check_markdown(text, ["Quarterly Operations Report"], ["North", "100", "120", "140"]))Run it for each converter on the same file and you have your own comparison.
When to use which
- Clean, single-column text, where speed matters more than structure: MarkItDown.
- Tables, scans and multi-column layouts, if you can afford the install: Docling.
- Marker, if you have a GPU or the llama.cpp dependency in place and you have read the model license. Test it on long documents first.
Another way to do it: anyformat
anyformat is a document extraction platform, and parsing is its first step, available on its own. You send a PDF and get markdown back from the API, with no models to install and no GPU.
import os
import time
from anyformat.sdk import Client
from anyformat.sdk.errors import StillParsing
client = Client(api_key=os.environ["ANYFORMAT_API_KEY"])
run = client.parse("report.pdf")
while True:
try:
markdown = client.get_markdown(run.id)
break
except StillParsing:
time.sleep(3)Install it with pip install anyformat. The markdown that comes back is structured: it contains anchors that identify each block of the page. The full parse result also lists those blocks with their page, bounding box and a confidence score, which is how you keep the link from a piece of text back to where it sits on the page. This example fetches only the markdown string; the block output is described in the API docs. The response shape is in the API docs. The same parse is reachable from an agent through the MCP server.
We did not run anyformat in the test above, so this page makes no claim about how its output compares on these files. Run the check from the section above on the markdown it returns and compare for yourself, using a fictional document, since real files can contain client data. If your documents are clean and you only need text for an index, a local open-source converter is the simpler choice.
Frequently asked questions
What is the best PDF to markdown converter for RAG?
It depends on your documents. On our test files Docling gave the most complete output, MarkItDown was the fastest and only reliable on clean text, and Marker needed extra setup for scans. Test on your own PDFs with a check like the one above before you pick.
Does MarkItDown do OCR?
Not for a plain PDF conversion. Its README describes a separate markitdown-ocr plugin that uses an LLM with vision to read images, and an option that uses Azure Document Intelligence.
Is Marker free for commercial use?
Its code is Apache 2.0. The model weights use a modified AI Pubs Open Rail-M license that its README describes as free for research, personal use and startups under $5M funding or revenue. Read the license in its repository for your case.
Why do tables break when converting a PDF to markdown?
A PDF stores characters at positions on a page, not rows and columns. A converter that only reads the text in order turns each cell into its own line. One that detects the table structure can rebuild the rows.
Do I need a GPU?
Not for the three tools we ran: all three converted text PDFs on a Mac without a GPU. Marker's README describes a GPU mode with higher accuracy, and its scanned-page mode on a Mac needed an extra binary, as described above.







