Extracting structured data from a PDF with an LLM means reading the document's text, describing the fields you want as a schema, and asking the model to return values that fit that schema as JSON. It works on layouts the model has never seen, which is why it replaced per-supplier templates. It also fails in predictable ways, and most of the work in production is catching those failures.
Most people do this for one of four reasons: to build a RAG over their own documents, to load the values into a database or spreadsheet, to trigger a workflow when a document arrives (an invoice to approve, a form to route), or to give an agent facts to work with instead of pages to read. In every case the point is the same: turn a document into fields a program can use.
This page does it end to end in Python: one working script, then the five places it breaks, then what to add so you can trust the output. The example is an invoice because everyone has one, but nothing below is specific to invoices.
What you will need
- Python 3.10 or newer.
- An API key for the LLM you want to call. The example uses OpenAI; swap in Anthropic or Gemini if you prefer.
- One PDF, ideally a fictional or sample invoice, because the examples print results and you should not send real client documents to logs or to a model you have not approved. Use the most complex one you have: a multi-page document with a table, a scan, or two columns. A clean one-page invoice makes every approach look good and hides the failures this page is about.
- Five minutes to install the libraries:
pip install pymupdf pydantic openaiStep 1. Get the text out of the PDF
A PDF is a drawing instruction set, not a text file, so the first step is turning pages into text the model can read. If the PDF has a text layer, a library such as PyMuPDF reads it directly.
import pymupdf
def pdf_to_text(path: str) -> str:
doc = pymupdf.open(path)
pages = [f"--- page {i + 1} ---\n{page.get_text()}" for i, page in enumerate(doc)]
return "\n".join(pages)Keeping the page markers matters later: they let you point at where a value came from. If this returns almost nothing, the PDF is a scan and has no text layer. That is failure 1 below.
Step 2. Describe what you want as a schema
A schema is a typed description of the output: field names, types and what each field means. Declaring it forces the model to return a value of the right shape instead of writing prose.
| Field | Type | Description the model reads |
|---|---|---|
| invoice_number | text | Invoice ID as printed, usually near the top |
| issued_on | date | Issue date, not the due date |
| supplier_name | text | Who issued the invoice |
| currency | text | ISO code, for example EUR or USD |
| purchase_order | text | Purchase order number, only if one is printed |
| tax | number | Total tax amount, only if shown separately |
| total | number | Grand total including tax |
| line_items | list | One entry per row in the table: description, quantity, unit price, amount before tax |
The same thing written as code with Pydantic, which is what you paste into the script:
from datetime import date
from pydantic import BaseModel, Field
class LineItem(BaseModel):
description: str
quantity: float
unit_price: float
amount: float = Field(description="Line amount before tax")
class Invoice(BaseModel):
invoice_number: str | None = Field(default=None, description="Invoice ID as printed, usually near the top")
issued_on: date | None = Field(default=None, description="Issue date, not the due date")
supplier_name: str | None = Field(default=None, description="Who issued the invoice")
currency: str | None = Field(default=None, description="ISO 4217 code, e.g. EUR, USD")
purchase_order: str | None = Field(default=None, description="Purchase order number, only if one is printed")
line_items: list[LineItem] = Field(default_factory=list)
tax: float | None = Field(default=None, description="Total tax amount, only if shown separately")
total: float | None = Field(default=None, description="Grand total including tax")The descriptions are the prompt. "Issue date, not the due date" prevents the most common mix-up on invoices, and it costs nothing. Every field is optional on purpose: if the document has no issue date or no total, the model must be able to say so instead of inventing one.
Step 3. Ask the model for JSON that fits the schema
Structured output means the model is constrained to return JSON that validates against your schema, so you get an object back, not text you have to parse. Most providers support it; OpenAI's structured outputs guide describes the OpenAI version. This example uses the OpenAI Python SDK, and the same pattern works with Anthropic or Gemini.
from openai import OpenAI
client = OpenAI() # reads OPENAI_API_KEY from the environment
def extract_invoice(text: str) -> Invoice:
response = client.responses.parse(
model="gpt-5", # use the current model id on the day you run this
input=[
{"role": "system", "content": (
"Extract the invoice fields from the document. "
"If a field is not present in the document, do not guess it."
)},
{"role": "user", "content": text},
],
text_format=Invoice,
)
return response.output_parsedThe function takes the text as an argument because the next step needs the same text to check the result against.
Step 4. Check the output before you trust it
A validation step is deterministic code that tests whether the extracted values are consistent with each other and with the source. It catches the errors that look right.
def check(invoice: Invoice, source_text: str) -> list[str]:
problems = []
# 1. Completeness: the fields you cannot do without must be present.
for name in ("invoice_number", "issued_on", "total"):
if getattr(invoice, name) is None:
problems.append(f"{name} is missing")
# 2. Arithmetic: with tax stated, lines plus tax must equal the total, to the cent.
# Without it, lines can only be checked against the total as an upper bound.
if invoice.total is not None:
lines_sum = round(sum(li.amount for li in invoice.line_items), 2)
if invoice.tax is not None:
if abs(lines_sum + invoice.tax - invoice.total) > 0.01:
problems.append("line items plus tax do not match the total")
elif lines_sum > invoice.total + 0.01:
problems.append("line items exceed the total")
# 3. Grounding: every identifier the model returns must appear in the document.
for name in ("invoice_number", "supplier_name", "purchase_order"):
value = getattr(invoice, name)
if value and value not in source_text:
problems.append(f"{name} not found in source text")
# 4. Plausibility: dates in the future are usually a misread.
if invoice.issued_on and invoice.issued_on > date.today():
problems.append("issued_on is in the future")
return problemssource_text = pdf_to_text("invoice.pdf")
invoice = extract_invoice(source_text)
problems = check(invoice, source_text)
print(problems or "all checks passed")The blocks on this page are one script: paste them into a single file in order and they run together. That is the whole pattern. This run prints the check results and not the invoice, and the walkthrough assumes a fictional invoice: with real documents, log field names and check results rather than extracted values, which can contain client data.
Here is what the checks are for. The values below are illustrative, not from a real document.
A good output
A good output passes every check. The lines add up, the number is on the page and the date is sensible:
{"invoice_number": "INV-2041", "issued_on": "2026-03-14",
"supplier_name": "Acme Hosting S.L.", "currency": "EUR",
"line_items": [{"description": "Hosting", "quantity": 1, "unit_price": 400.0, "amount": 400.0},
{"description": "Support", "quantity": 2, "unit_price": 150.0, "amount": 300.0}],
"tax": 147.0, "total": 847.0}check() returns an empty list. The lines (700) plus the tax field (147) equal the total (847). The arithmetic check is exact only when the invoice states its tax separately. When it does not, the check can only tell that the lines exceed the total, so an under-read line goes unnoticed; that is why the other checks matter.
A bad output
A bad output looks just as tidy, and that is the danger. The model misread the table and invented a number:
{"invoice_number": "INV-2014", "issued_on": "2031-03-14",
"supplier_name": "Acme Hosting S.L.", "currency": "EUR",
"line_items": [{"description": "Hosting", "quantity": 1, "unit_price": 400.0, "amount": 400.0},
{"description": "Support", "quantity": 2, "unit_price": 1500.0, "amount": 3000.0}],
"tax": 147.0, "total": 847.0}check() returns three problems: line items plus tax do not match the total, invoice_number is not in the source text and issued_on is in the future. Nothing in the JSON itself said any of this was wrong; only the checks did.
The grounding check is the cheap version of a large idea: a value that cannot be traced back to the page should not survive your pipeline. The longer treatment of that idea is in how to reduce LLM hallucinations in document extraction.
Where the script breaks
Failure 1: scanned PDFs
A scan has no text layer, so pdf_to_text returns almost nothing and the model is asked to extract from empty text. Detect it (less than a few hundred characters of text per page, not counting the page markers) and route the file to OCR or to a vision model that reads the page image. Never send the empty text to the LLM: it will often return plausible values anyway.
Failure 2: tables
page.get_text() flattens a table into a stream where columns and rows can interleave. A line item's quantity lands next to the wrong price and the arithmetic check in step 4 fails. Extract tables as tables (a layout-aware parser, or a vision model), or pass the model page images for the pages that contain them.
Failure 3: long documents
A 200-page contract does not fit comfortably in one prompt, and models read the middle of a long context less reliably than the start and end. Split by section or page range, extract per chunk, then merge. The merge step is where duplicates and conflicts show up, so keep the page markers. There is more on this in long-document extraction.
Failure 4: fields that are not in the document
Ask for purchase_order on an invoice without one and a model may return a plausible number instead of nothing. Make absence a legal answer: declare every field optional, as in the schema above, tell the model explicitly to return null rather than guess (the system prompt does), and let the grounding check catch the rest: it searches the document for the invoice number, the supplier and the purchase order, so an invented one is reported.
Failure 5: no way to tell right from wrong at scale
The script returns the same clean JSON whether it read the page correctly or not. If you process ten documents you can eyeball them. If you process ten thousand you need a signal per field that says which ones to review. Raw model confidence is a poor signal, because it tends to be highest where the model is filling a gap. This is the problem the next section is about.
What to add for production
Three things separate a script from a pipeline: a pointer from each value back to where it was read, a confidence score you can set a threshold on, and a place where a person reviews what falls below it. You can build all three on top of the code above. The grounding check is a start, a calibration set and a threshold come next, and a review interface is the largest piece.
Another way to do it: anyformat
anyformat is a document extraction platform, and the script above is the job it does for you. The steps are the same (parse the document, apply a schema, return fields), but the pointer to the source, the confidence score and the review interface are already built.
The workflow is defined once as a typed graph, a parse node feeding an extract node, then each document is one upload and one poll. Each field comes back with its value, a confidence score and the evidence it was read from.
import os
from anyformat.sdk import Client
from anyformat.workflow import Schema
client = Client(api_key=os.environ["ANYFORMAT_API_KEY"])
result = (
client.workflow("Invoice")
.parse()
.extract([
Schema.string("invoice_number", "Invoice ID as printed, usually near the top."),
Schema.date("issued_on", "Issue date, not the due date."),
Schema.float("total", "Grand total including tax."),
])
.create()
.run("invoice.pdf")
.wait()
)
field = result.fields["total"]
print(field.value, field.confidence) # value is a string; confidence runs from 0 to 100
for e in field.evidence:
print(e.page_number, e.text)Install it with pip install anyformat. As above, use a fictional document while you try it: the example prints a value and its evidence, and real documents can contain client data. A field the model could not find comes back as null, with a low confidence score so you can filter it. The evidence is the source text and the page number, so a reviewer checks a flagged field against the page in seconds. The response shape and the full node list are in the API docs. The reasoning behind per-field confidence is in calibrated confidence, and the same workflow is reachable from an agent through the MCP server.
anyformat is one option among several. If you process a few dozen documents of one type, the script above plus the checks in step 4 is enough, and you should not pay for more. The case for a platform starts when volume, variety or the cost of a wrong value makes the review layer the real project.
Frequently asked questions
Can an LLM extract data from a PDF directly?
Yes, if you give it the PDF's text or page images and a schema for the output. The hard part is not the extraction call. It is handling scans, tables and long documents, and knowing which of the returned values to trust.
What is the difference between structured output and JSON mode?
JSON mode guarantees the reply is valid JSON. Structured output guarantees it also matches your schema: the field names, the types and the required fields. For extraction you want the second.
How do I stop an LLM from making up values in extraction?
You cannot bring it to zero, so make it detectable: declare fields optional, instruct the model to return null instead of guessing, verify that each value appears in the source, and validate values against each other. More on this in reducing hallucinations in document extraction.
Does this work on scanned PDFs?
Not with a text library alone. A scan has no text layer, so you need OCR or a vision model first. Detect the empty text layer and route accordingly.
How much does LLM extraction cost per page?
It depends on the model, the page length and whether you send text or images. Measure it on 20 of your own documents before you pick a model, because public price lists change often enough that any number here would go stale.







