How PDF table extraction actually works
No machine learning, no ruling-line detection, no training set. Tables are found by measuring horizontal gaps and vertical bands, which is why the same file returns the same bytes every time.
Step one: segments
Runs are grouped into visual lines by baseline, then each line is split wherever the horizontal gap exceeds 0.7 of the font size. The pieces are called segments, and a segment is a candidate cell.
This handles the case that trips up most approaches: right-aligned numeric columns. "10" and "50,000.00" have wildly different left edges, so clustering cells by left edge loses precisely the money columns people care about. Splitting by gap does not care where a cell starts.
Step two: columns, not counts
The obvious rule is that a table is a run of lines with the same number of segments. It is also wrong, and wrong in a way that only real documents reveal: an empty cell produces no segment, so a row with a blank in it has fewer segments and ends the table.
That is not an edge case. It is the remittance deduction column, the vacant unit’s lease dates, the timesheet weekend, the lab result with no flag — every optional field on every form. Measured against twelve realistic document shapes, this single rule broke four of them, and quietly: a seven-row lab report came back with three rows and a 200 OK.
So the columns are worked out first. The widest lines in a run define a set of x bands; every other line is then tested against those bands. A line whose segments each land in a distinct band, left to right, is a row of that table — with blanks wherever a band is empty.
| Rule | What it gets wrong |
|---|---|
| Same segment count | Any row with an empty cell ends the table |
| Cluster by left edge | Right-aligned money columns scatter |
| Ruling lines only | Borderless tables are invisible |
| Column bands, widest-first | Merged header cells — reported, not guessed |
Step three: the awkward details
Bands are measured from whichever lines carry every column, which can be very few. A band measured from three rows of "929.00" is narrower than the "1,393.50" on the fourth — so a cell is matched to the band it OVERLAPS most, not the band that contains it.
A cell that overlaps nothing — a short right-aligned label sitting in a gutter — is assigned to the nearest band by centre rather than rejected. Rejecting it made the whole line unmappable, and an invoice’s "Subtotal / Tax / Total" block vanished from the response entirely.
A run of consecutive sparse lines is its own table rather than the tail of the one above. Three two-column rows under a four-column table are a totals block, and returning them as blank-padded rows of the line items is wrong in a way anyone summing a column would notice.
A single-segment line — a heading or a caption — ends a run. Each table then gets its own columns, which matters because two tables under separate headings were otherwise measured against each other’s bands, and the narrower one’s header did not fit and was dropped.
What it still cannot do
A header merged across columns with colspan is one segment over three. It is not part of the block, and no amount of band matching changes that. The response reports no header and a confidence of zero rather than inventing a split — which is the honest answer, if not the useful one.
Two tables printed side by side with a narrow gutter can merge into one wide table. The gutter between them is the same shape as a column gap, and nothing in the geometry distinguishes the two.
See it on a real document
The extraction playground takes a PDF and shows the JSON with every bounding box drawn over the page. No key needed for the first few.
Related reading
- How a PDF stores text, and why that makes extraction hard
- How a table header is detected in a PDF
- Why this API does not do OCR, and how to tell if you need it
- PDF coordinates and bounding boxes, explained
- Deterministic PDF extraction versus an LLM
- What the confidence scores mean
- What it takes to run headless Chromium in production