PDFCraft
How it works

How a table header is detected in a PDF

A header row looks exactly like a data row. Nothing in the file marks it. Here is the evidence that is actually available, what each piece is worth, and why the answer is a number rather than a yes.

The rule that was not enough

The obvious signal is numeric contrast: a column that parses as a number in every body row, under a cell in that column that does not. "Quantity" over 10, 2, 2, 1.

It is a good signal and it fires on invoices, statements and anything financial. It also cannot fire at all on a table with no numeric column — a schedule, a roster, a parts list — which is a great many tables. Using it alone meant `header` came back null on almost everything, and a caller had no way to tell that from "row 0 really is data".

What else is available

Five more signals, each independent of the others, each weighted by how much it is actually worth:

SignalWeightWhy that much
Numeric contrast0.55Enough on its own. A column that is a number in every row and a word in the first is not a coincidence.
Set in a face the body never uses0.40Strong, but not alone — one italic cell should not invent a header.
Larger than the body0.20Corroborating. Headers are often the same size in bold.
The same row atop the next page0.50Near proof. Data does not reappear verbatim after a page break.
Text appears nowhere in the body0.10Weak. Plenty of first data rows are unique too.
No digits in the row, digits below0.10Weak, same reason.
Mostly blank row−0.40A spacer, not a header.

Why repetition had to be gathered first

The strongest signal is the cheapest to miss. A long statement is one table the renderer broke across fourteen pages, repeating the header each time — and a row that reappears verbatim at the top of the next page is a header almost by definition.

The original stitcher required both tables to already HAVE a header before it would merge them, which is circular: a statement whose header numeric contrast could not detect never stitched, so it never earned the repetition evidence that would have detected it. Comparing the candidate row instead, and feeding the result back into the score, breaks the loop.

Why you get a number

Promotion happens at 0.5, and the score comes back on every table as `header_confidence`. That is deliberate. A caller writing `row[headers.indexOf("Amount")]` deserves to know whether that lookup rests on a numeric column that could not be anything else, or on a row that merely happened to be in a different font.

The weights are judgement rather than calibration — there is no labelled corpus behind them, and saying otherwise would be the dishonest part. They live in one named table in the source so they can be argued with.

{
  "header": ["Reference", "Counterparty", "Amount"],
  "header_confidence": 1,
  "confidence": 0.95,
  "rows": [["TX-2401", "Caldera Print", "137.25"]]
}

See it on a real document

The extraction playground takes a PDF and shows the JSON with every bounding box drawn over the page. No key needed for the first few.

Related reading