PDFCraft
How it works

How PDF table extraction actually works

No machine learning, no ruling-line detection, no training set. Tables are found by measuring horizontal gaps and vertical bands, which is why the same file returns the same bytes every time.

Step one: segments

Runs are grouped into visual lines by baseline, then each line is split wherever the horizontal gap exceeds 0.7 of the font size. The pieces are called segments, and a segment is a candidate cell.

This handles the case that trips up most approaches: right-aligned numeric columns. "10" and "50,000.00" have wildly different left edges, so clustering cells by left edge loses precisely the money columns people care about. Splitting by gap does not care where a cell starts.

Step two: columns, not counts

The obvious rule is that a table is a run of lines with the same number of segments. It is also wrong, and wrong in a way that only real documents reveal: an empty cell produces no segment, so a row with a blank in it has fewer segments and ends the table.

That is not an edge case. It is the remittance deduction column, the vacant unit’s lease dates, the timesheet weekend, the lab result with no flag — every optional field on every form. Measured against twelve realistic document shapes, this single rule broke four of them, and quietly: a seven-row lab report came back with three rows and a 200 OK.

So the columns are worked out first. The widest lines in a run define a set of x bands; every other line is then tested against those bands. A line whose segments each land in a distinct band, left to right, is a row of that table — with blanks wherever a band is empty.

RuleWhat it gets wrong
Same segment countAny row with an empty cell ends the table
Cluster by left edgeRight-aligned money columns scatter
Ruling lines onlyBorderless tables are invisible
Column bands, widest-firstMerged header cells — reported, not guessed

Step three: the awkward details

Bands are measured from whichever lines carry every column, which can be very few. A band measured from three rows of "929.00" is narrower than the "1,393.50" on the fourth — so a cell is matched to the band it OVERLAPS most, not the band that contains it.

A cell that overlaps nothing — a short right-aligned label sitting in a gutter — is assigned to the nearest band by centre rather than rejected. Rejecting it made the whole line unmappable, and an invoice’s "Subtotal / Tax / Total" block vanished from the response entirely.

A run of consecutive sparse lines is its own table rather than the tail of the one above. Three two-column rows under a four-column table are a totals block, and returning them as blank-padded rows of the line items is wrong in a way anyone summing a column would notice.

A single-segment line — a heading or a caption — ends a run. Each table then gets its own columns, which matters because two tables under separate headings were otherwise measured against each other’s bands, and the narrower one’s header did not fit and was dropped.

What it still cannot do

A header merged across columns with colspan is one segment over three. It is not part of the block, and no amount of band matching changes that. The response reports no header and a confidence of zero rather than inventing a split — which is the honest answer, if not the useful one.

Two tables printed side by side with a narrow gutter can merge into one wide table. The gutter between them is the same shape as a column gap, and nothing in the geometry distinguishes the two.

See it on a real document

The extraction playground takes a PDF and shows the JSON with every bounding box drawn over the page. No key needed for the first few.

Related reading