Extracting a table from a PDF means getting its rows and columns back as data, a list of rows or a pandas DataFrame, instead of a pile of text. It is harder than it looks, because a PDF stores characters at positions on a page and has no concept of a row or a column. A library has to infer the table from the lines drawn around the cells or from how the words line up.
You usually want this to load figures into a spreadsheet or database, to feed a table into an analysis, or to give a language model a table it can read. This page runs the three most used Python libraries on the same five fictional PDFs, shows the code that worked, and reports where each one broke.
What you will need
- Python 3.10 or newer.
- One PDF with a table in it, ideally a fictional or sample document, because the examples print what they find and text from real client documents can end up in terminal and job logs. Use the hardest one you have: no ruling lines, merged header cells or a table that runs across pages. A tidy ruled table makes every library look good.
- One virtual environment per library, so their dependencies do not collide:
pip install pdfplumber # environment 1
pip install camelot-py # environment 2
pip install pymupdf # environment 3Step 1. Check that the PDF has a text layer
All three libraries read the text stored in the PDF. A scan is an image with no text layer, and each of them returned an empty list on our scanned test file without raising an error. Check every page first, because a text PDF can have a short cover page or a blank one:
import pdfplumber
with pdfplumber.open("table.pdf") as pdf:
pages_without_text = [
number
for number, page in enumerate(pdf.pages, start=1)
if len((page.extract_text() or "").strip()) <= 50
]
if pages_without_text:
print("pages with little or no text (possible scans):", pages_without_text)
else:
print("every page has a text layer")A page listed there is worth a look rather than a verdict, since a cover or a blank page also has little text. If a page with a table has no text layer, none of the code below will work and you need OCR or a layout-aware parser first; see the best Python OCR libraries for the OCR options.
Step 2. Extract the table with each library
pdfplumber, the lightest of the three, reads ruled tables by default:
import pdfplumber
with pdfplumber.open("table.pdf") as pdf:
tables = pdf.pages[0].extract_tables()
if tables:
print(tables[0]) # a list of rows
else:
print("no table found")Camelot returns a pandas DataFrame, with lattice for ruled tables and stream for borderless ones:
import camelot
tables = camelot.read_pdf("table.pdf", pages="all", flavor="lattice")
print(len(tables))
if len(tables):
print(tables[0].df)PyMuPDF finds tables with find_tables, and the result converts to a DataFrame too:
import pymupdf
doc = pymupdf.open("table.pdf")
tabs = doc[0].find_tables()
if tabs.tables:
print(tabs[0].extract()) # a list of rows
df = tabs[0].to_pandas()What we saw
We generated five fictional PDFs with a known table: one with ruling lines, one without any, one with merged header cells and a multi-line cell, one that runs across two pages with a repeated header, and one that is a scan. We compared every output cell by cell against the table we generated. The versions were pdfplumber 0.11.10, Camelot 2.0.0 and PyMuPDF 1.28.2, and each library was run in its default mode and in its text-based mode.
| pdfplumber | Camelot | PyMuPDF | |
|---|---|---|---|
| Ruled table | Correct | Correct | Correct |
| No ruling lines, default mode | Empty | Empty | Empty |
| No ruling lines, text mode | Right cells plus blank rows | Right cells plus a title row | Right cells plus blank rows |
| Merged header, multi-line cell | Correct in default mode; broken in text mode | Correct in lattice; broken in stream | Correct in default mode; broken in text mode |
| Two pages | Two tables, header repeated | Two tables, header repeated | Two tables, header repeated |
| Scan, no text layer | Empty list, no warning | Empty list, no warning | Empty list, no warning |
This is five small, clean, generated files, so it shows how each library fails and says nothing about how often it fails on your documents.
Ruling lines decide everything in the default modes
The default mode of each library is the same idea with a different name: find the lines drawn around the cells and read the text inside them. On the ruled table all three were correct. On the table with no ruling lines all three returned nothing, and on that one you have to switch to the text-based strategy, which guesses cells from word positions.
The text-based mode works but needs cleaning
With vertical_strategy and horizontal_strategy set to "text", pdfplumber found the right cells but added a blank row between every row of the table (11 rows where 6 were expected):
import pdfplumber
with pdfplumber.open("table.pdf") as pdf:
found = pdf.pages[0].extract_tables(
{"vertical_strategy": "text", "horizontal_strategy": "text"}
)
rows = [r for r in found[0] if any(r)] if found else [] # drop the blank rowsAfter dropping empty rows the table was correct. PyMuPDF's text strategy gave the same blank rows, and Camelot's stream mode added the page title as a first row. On the table with merged header cells and a multi-line cell, the text modes broke: the cell text was split across rows and only one of five rows came out intact in pdfplumber and PyMuPDF. PyMuPDF also cut "reorder monthly" to "reorder mont" in that run.
Merged cells and page breaks are not handled for you
A merged header such as "Sales" over two sub-columns comes out as "Sales" followed by an empty cell, with the sub-headers on a second row. A table that runs across two pages comes back as two tables, each with its own copy of the header, and none of the three libraries joins them. You concatenate the pieces and drop the repeated header yourself.
A scan returns an empty list, not an error
On the scanned file every library returned an empty result without a warning, which is easy to miss in a pipeline. That is why Step 1 checks for a text layer first.
The same test on a real financial PDF
The video runs the three libraries on a real financial document from the public ParseBench dataset: two tables with 45 companies, multi-line names and merged cells. Every row was checked against the ground truth. The results shown in the video were:
| Approach | Rows correct | What went wrong |
|---|---|---|
| pdfplumber | 51.1% | Drops names in merged cells |
| Camelot, lattice | 0% | No ruling lines to detect |
| Camelot, stream | 88.9% | Merges the two tables |
| PyMuPDF | 82.2% | Scrambles text in complex cells |
| anyformat | 95.6% | 43 of 45 companies |
It is one document and one run each, so read it as a pattern and not as a benchmark. The pattern matches the synthetic test above: ruling lines decide the default modes, and the text-based modes keep more rows but damage them.
None of the three libraries tells you that a row is wrong. They return a table either way, so a damaged cell looks exactly like a correct one, and that is why the check below is worth writing. anyformat returns the same table with a confidence score per field and the source text each value was read from, which is what lets a person check only the doubtful rows.
Check on your own PDFs
Write down one row you know is in your table and test for it. The function reports only which check failed, never the cell contents, so its output is safe to log.
def check_table(table: list[list[str]], expected_row: list[str]) -> list[str]:
problems = []
if not table:
problems.append("no table found")
elif expected_row not in table:
problems.append("expected row is not intact")
return problems
print(check_table(tables[0] if tables else [], ["North", "1200", "1350", "1480"]))Run it for each library on the same file and you have a comparison for your documents, which is worth more than ours.
Reproduce the test
The five PDFs come from a short script, so you can rerun everything. It needs pip install reportlab:
import json
from reportlab.lib.pagesizes import A4
from reportlab.platypus import Table, TableStyle, SimpleDocTemplate, Paragraph, Spacer
from reportlab.lib import colors
from reportlab.lib.styles import getSampleStyleSheet
hdr = ["Region", "Q1", "Q2", "Q3"]
rows = [["North", "1200", "1350", "1480"], ["South", "980", "1010", "1100"],
["East", "1500", "1620", "1710"], ["West", "760", "810", "905"],
["Total", "4440", "4790", "5195"]]
def simple(filename, data, grid):
style = [("FONTSIZE", (0, 0), (-1, -1), 11), ("ALIGN", (1, 0), (-1, -1), "RIGHT"),
("BOTTOMPADDING", (0, 0), (-1, -1), 6), ("TOPPADDING", (0, 0), (-1, -1), 6)]
if grid:
style += [("GRID", (0, 0), (-1, -1), 0.8, colors.black),
("BACKGROUND", (0, 0), (-1, 0), colors.lightgrey)]
table = Table(data, colWidths=[110, 90, 90, 90])
table.setStyle(TableStyle(style))
SimpleDocTemplate(filename, pagesize=A4).build(
[Paragraph("Sales by region", getSampleStyleSheet()["Heading2"]), Spacer(1, 12), table])
simple("a_ruled.pdf", [hdr] + rows, grid=True)
simple("b_borderless.pdf", [hdr] + rows, grid=False)
merged = [["Product", "Sales", "", " Notes"], ["", "H1", "H2", ""],
["Widget", "100", "120", "Best seller\nreorder monthly"],
["Gadget", "80", "95", "Discontinued\nin Q4"], ["Gizmo", "60", "70", "New"]]
table = Table(merged, colWidths=[90, 70, 70, 150])
table.setStyle(TableStyle([("GRID", (0, 0), (-1, -1), 0.8, colors.black),
("SPAN", (1, 0), (2, 0)), ("SPAN", (0, 0), (0, 1)), ("SPAN", (3, 0), (3, 1)),
("ALIGN", (0, 0), (-1, 1), "CENTER"), ("VALIGN", (0, 0), (-1, -1), "MIDDLE")]))
SimpleDocTemplate("c_merged.pdf", pagesize=A4).build([table])
long = [hdr] + [[f"Item{i:02d}", str(100 + i * 7), str(200 + i * 3), str(300 + i * 11)] for i in range(1, 61)]
table = Table(long, colWidths=[110, 90, 90, 90], repeatRows=1)
table.setStyle(TableStyle([("GRID", (0, 0), (-1, -1), 0.8, colors.black),
("BACKGROUND", (0, 0), (-1, 0), colors.lightgrey), ("FONTSIZE", (0, 0), (-1, -1), 11)]))
SimpleDocTemplate("d_twopage.pdf", pagesize=A4).build([table])For the scan, render the ruled PDF to an image and save it back as a PDF with no text layer:
import pymupdf
src = pymupdf.open("a_ruled.pdf")
pix = src[0].get_pixmap(dpi=150)
out = pymupdf.open()
page = out.new_page(width=pix.width, height=pix.height)
page.insert_image(page.rect, pixmap=pix)
out.save("e_scan.pdf")Install and licence notes
Camelot is the heaviest install of the three, around 157 MB with OpenCV, against about 42 MB for pdfplumber and 53 MB for PyMuPDF in our environments. Camelot 2.0.0 no longer needs Ghostscript by default and converts pages through pypdfium2. Older 0.x versions, which most tutorials still cover, failed on every default call with "Ghostscript is not installed" in our run, even with the gs binary on the path, so install the current version.
pdfplumber is MIT licensed and Camelot is MIT licensed. PyMuPDF is dual licensed under AGPL 3.0 or a commercial license from Artifex, which matters if you ship closed-source software; read its licensing page for your case. All three had a release within the last four months on 2026-10-01.
When these libraries are not enough: anyformat
These libraries fit PDFs with a text layer and tables that are ruled or cleanly aligned. They stop fitting when tables have no lines, merged cells, multi-line cells or page breaks, or when the files are scans. For those cases people turn to a layout-aware parser, a vision language model or an OCR step that finds table structure.
anyformat's parsing is one option. Its parse result lists a document's blocks, and table blocks carry their rows, according to the SDK reference. We did not run it on the five synthetic files; in the video it ran on the real financial PDF above and got 43 of 45 companies right, which is one document and one run. For long tables specifically, see long-document extraction.
Watch the test
The video version of this test runs about five minutes and covers the three libraries and the real-PDF results above.
Frequently asked questions
What is the best Python library to extract tables from a PDF?
It depends on the table. For ruled tables in text-based PDFs all three worked in our test, and pdfplumber was the lightest to install. For tables without ruling lines, none of the defaults worked and every library needed its text-based mode plus cleanup.
Can pdfplumber extract tables from a scanned PDF?
No. It reads the text layer, and a scan has none, so it returns an empty list. You need OCR or a layout-aware parser first.
How do I get a PDF table into a pandas DataFrame?
Camelot returns one directly (tables[0].df), PyMuPDF converts with to_pandas(), and with pdfplumber you pass the list of rows to pandas.DataFrame, using the first row as the header.
Why does my table come out with blank rows?
The text-based strategies in pdfplumber and PyMuPDF find the cell edges from word positions and can add empty rows between real ones. Drop the rows where every cell is empty.
How do I extract a table that spans several pages?
Each library returns one table per page. Concatenate the tables and drop the repeated header rows yourself.







