What the confidence scores mean
Every field and every table comes back with a number between 0 and 1. Here is exactly what goes into each one, so you can set a threshold on evidence rather than on vibes.
Field confidence
How much of a label/value pairing was read off the page, and how much was inferred.
| Evidence | Effect |
|---|---|
| Base | 0.50 |
| Label and value in one segment — "Total: 1,200.00" | +0.20 |
| Short value (under 60 characters) | +0.10 |
| The type you asked for parsed cleanly | +0.10 |
| Value reassembled across a line wrap | −0.20 |
| Very long value (over 120 characters) | −0.20 |
| You asked for a number and got something that is not one | −0.15 |
Why inline scores higher
An inline "Total: 1,200.00" is the document’s own punctuation saying those two things belong together. A value taken from the next segment along is a judgement — it might be the next column of a table rather than the value of that label.
A value reassembled across a wrap involves more judgement still: correct far more often than not, but it is the mechanism that could swallow the block underneath.
Table confidence
Separate from `header_confidence`, and asking a different question: is this block a table at all, or a layout artifact? A page is a grid of boxes and plenty of things in a grid are not tables — a two-column form, an address beside a logo, a row of footer links.
| Evidence | Effect |
|---|---|
| Base (it already cleared the geometry gate) | 0.45 |
| Six or more rows | +0.15 |
| Fifteen or more rows | +0.10 |
| A header was promoted | +0.15 |
| A column is numeric in most rows | +0.15 |
| Columns tightly aligned | +0.10 |
| Continued onto the next page | +0.15 |
| A cell holding a paragraph | −0.30 |
| Two thin columns, few rows, no numbers | −0.25 |
| More empty cells than full ones | −0.15 |
Penalties are modifiers, not vetoes
A real table with one long description cell is still a table, and a penalty does not throw it away. What a penalty does is tip a block that had little going for it in the first place below the bar.
The default bar is 0.5, which keeps everything the geometry gate lets through unless something argues against it. `options.min_confidence` moves it, and `tables_suppressed` always reports how many blocks fell below whatever bar was in force — so a bar set too high is visible rather than silent.
{
"tables": [ /* … */ ],
"tables_suppressed": 3
}These are judgement, not calibration
There is no labelled corpus behind these weights. They are engineering judgement, written down in one named table in the source so they can be argued with and tuned rather than buried inside an expression.
Saying otherwise would be the dishonest part. Use the numbers as an ordering — this answer rests on more evidence than that one — rather than as a probability.
See it on a real document
The extraction playground takes a PDF and shows the JSON with every bounding box drawn over the page. No key needed for the first few.
Related reading
- How a PDF stores text, and why that makes extraction hard
- How PDF table extraction actually works
- How a table header is detected in a PDF
- Why this API does not do OCR, and how to tell if you need it
- PDF coordinates and bounding boxes, explained
- Deterministic PDF extraction versus an LLM
- What it takes to run headless Chromium in production