How a table header is detected in a PDF
A header row looks exactly like a data row. Nothing in the file marks it. Here is the evidence that is actually available, what each piece is worth, and why the answer is a number rather than a yes.
The rule that was not enough
The obvious signal is numeric contrast: a column that parses as a number in every body row, under a cell in that column that does not. "Quantity" over 10, 2, 2, 1.
It is a good signal and it fires on invoices, statements and anything financial. It also cannot fire at all on a table with no numeric column — a schedule, a roster, a parts list — which is a great many tables. Using it alone meant `header` came back null on almost everything, and a caller had no way to tell that from "row 0 really is data".
What else is available
Five more signals, each independent of the others, each weighted by how much it is actually worth:
| Signal | Weight | Why that much |
|---|---|---|
| Numeric contrast | 0.55 | Enough on its own. A column that is a number in every row and a word in the first is not a coincidence. |
| Set in a face the body never uses | 0.40 | Strong, but not alone — one italic cell should not invent a header. |
| Larger than the body | 0.20 | Corroborating. Headers are often the same size in bold. |
| The same row atop the next page | 0.50 | Near proof. Data does not reappear verbatim after a page break. |
| Text appears nowhere in the body | 0.10 | Weak. Plenty of first data rows are unique too. |
| No digits in the row, digits below | 0.10 | Weak, same reason. |
| Mostly blank row | −0.40 | A spacer, not a header. |
Why repetition had to be gathered first
The strongest signal is the cheapest to miss. A long statement is one table the renderer broke across fourteen pages, repeating the header each time — and a row that reappears verbatim at the top of the next page is a header almost by definition.
The original stitcher required both tables to already HAVE a header before it would merge them, which is circular: a statement whose header numeric contrast could not detect never stitched, so it never earned the repetition evidence that would have detected it. Comparing the candidate row instead, and feeding the result back into the score, breaks the loop.
Why you get a number
Promotion happens at 0.5, and the score comes back on every table as `header_confidence`. That is deliberate. A caller writing `row[headers.indexOf("Amount")]` deserves to know whether that lookup rests on a numeric column that could not be anything else, or on a row that merely happened to be in a different font.
The weights are judgement rather than calibration — there is no labelled corpus behind them, and saying otherwise would be the dishonest part. They live in one named table in the source so they can be argued with.
{
"header": ["Reference", "Counterparty", "Amount"],
"header_confidence": 1,
"confidence": 0.95,
"rows": [["TX-2401", "Caldera Print", "137.25"]]
}See it on a real document
The extraction playground takes a PDF and shows the JSON with every bounding box drawn over the page. No key needed for the first few.
Related reading
- How a PDF stores text, and why that makes extraction hard
- How PDF table extraction actually works
- Why this API does not do OCR, and how to tell if you need it
- PDF coordinates and bounding boxes, explained
- Deterministic PDF extraction versus an LLM
- What the confidence scores mean
- What it takes to run headless Chromium in production