Deterministic PDF extraction versus an LLM
Not a rivalry. They fail differently, cost differently by two orders of magnitude, and the sensible architecture uses both — with the cheap one first.
Where each one wins
| Geometry | Language model | |
|---|---|---|
| Native-text PDF | Milliseconds, exact | Seconds, usually right |
| Scanned page | Refuses, instantly | Works, with a vision model |
| Same file twice | Identical bytes | May differ |
| Unstated field | null | May invent one |
| Provenance | A box on the page | None — it never saw coordinates |
| Novel layout | May miss a table | Usually copes |
| Cost per 1,000 pages | Pennies | Dollars |
The failure modes are opposite, and that is the point
Geometry fails loudly. A table it cannot resolve comes back with a low confidence or not at all, and a page with no text layer returns a 422 naming the blank pages. You know.
A model fails quietly. Asked for a total that is not printed on the page, it will frequently produce a number — plausible, correctly formatted, and wrong. Nothing about the response says so.
For a finance pipeline that difference is the whole argument. A missing value is a retry. A confidently wrong value is a reconciliation problem discovered weeks later.
The architecture that actually makes sense
Try the deterministic path first. Take the answer when the confidence is high — which, on native-text business documents, is most of the time. Fall back to a model for the scans and the pages that came back uncertain.
You get the model’s coverage at a fraction of the cost and latency, and every value the fast path returned is one you can point at on the page. `min_confidence` and the per-field confidence scores exist so that routing decision can be made by your code rather than by eye.
const result = await extract({ file, schema, options: { min_confidence: 0.7 } });
const uncertain = Object.entries(result.fields)
.filter(([, field]) => field === null || field.confidence < 0.7)
.map(([name]) => name);
if (uncertain.length > 0) {
// Only these fields, only this page, only now.
await askTheModel(file, uncertain, result.pages_without_text);
}Where we are the wrong tool
Scans and photographs. Handwriting. A field that requires reading the document rather than locating it — "what is this contract about". Anything where the layout is genuinely novel every time and there is no structure to find.
Saying so is not modesty. Ranking for OCR intent would buy refund requests, and a page that hedges about it would be worse than useless to someone trying to choose.
See it on a real document
The extraction playground takes a PDF and shows the JSON with every bounding box drawn over the page. No key needed for the first few.
Related reading
- How a PDF stores text, and why that makes extraction hard
- How PDF table extraction actually works
- How a table header is detected in a PDF
- Why this API does not do OCR, and how to tell if you need it
- PDF coordinates and bounding boxes, explained
- What the confidence scores mean
- What it takes to run headless Chromium in production