PDFCraft
How it works

Deterministic PDF extraction versus an LLM

Not a rivalry. They fail differently, cost differently by two orders of magnitude, and the sensible architecture uses both — with the cheap one first.

Where each one wins

GeometryLanguage model
Native-text PDFMilliseconds, exactSeconds, usually right
Scanned pageRefuses, instantlyWorks, with a vision model
Same file twiceIdentical bytesMay differ
Unstated fieldnullMay invent one
ProvenanceA box on the pageNone — it never saw coordinates
Novel layoutMay miss a tableUsually copes
Cost per 1,000 pagesPenniesDollars

The failure modes are opposite, and that is the point

Geometry fails loudly. A table it cannot resolve comes back with a low confidence or not at all, and a page with no text layer returns a 422 naming the blank pages. You know.

A model fails quietly. Asked for a total that is not printed on the page, it will frequently produce a number — plausible, correctly formatted, and wrong. Nothing about the response says so.

For a finance pipeline that difference is the whole argument. A missing value is a retry. A confidently wrong value is a reconciliation problem discovered weeks later.

The architecture that actually makes sense

Try the deterministic path first. Take the answer when the confidence is high — which, on native-text business documents, is most of the time. Fall back to a model for the scans and the pages that came back uncertain.

You get the model’s coverage at a fraction of the cost and latency, and every value the fast path returned is one you can point at on the page. `min_confidence` and the per-field confidence scores exist so that routing decision can be made by your code rather than by eye.

const result = await extract({ file, schema, options: { min_confidence: 0.7 } });

const uncertain = Object.entries(result.fields)
  .filter(([, field]) => field === null || field.confidence < 0.7)
  .map(([name]) => name);

if (uncertain.length > 0) {
  // Only these fields, only this page, only now.
  await askTheModel(file, uncertain, result.pages_without_text);
}

Where we are the wrong tool

Scans and photographs. Handwriting. A field that requires reading the document rather than locating it — "what is this contract about". Anything where the layout is genuinely novel every time and there is no structure to find.

Saying so is not modesty. Ranking for OCR intent would buy refund requests, and a page that hedges about it would be worse than useless to someone trying to choose.

See it on a real document

The extraction playground takes a PDF and shows the JSON with every bounding box drawn over the page. No key needed for the first few.

Related reading