Why this API does not do OCR, and how to tell if you need it
Adding OCR would destroy both things that make this useful: the latency and the determinism. Here is how to tell in one request whether your documents need it.
The two kinds of PDF
A PDF generated by software has a text layer — the actual characters, with positions. Reading it is a geometry problem and it takes milliseconds.
A PDF produced by a scanner or a phone camera is a picture of a document. There are no characters in the file at all, only pixels. Getting text out means recognising shapes, which is a fundamentally different operation with a fundamentally different cost.
Most business documents are the first kind. Invoices from accounting systems, statements from banks, reports from BI tools, anything emailed rather than posted — all native text.
What OCR would cost
Roughly 30 ms a page becomes roughly 2 seconds a page — a factor of sixty. Deterministic output becomes probabilistic: run the same file twice and the numbers can differ. And a clear "this has no text layer, send it elsewhere" becomes a confident, plausible, wrong answer.
That last one is the real objection. A parser that returns nothing is annoying. A parser that returns 8 where the document said 3 is a liability, and you find out about it downstream.
How to tell which you have, in one request
Send the file. If there is no text layer you get a 422 with the code `extraction_failed` and, in the body, exactly which pages were blank. It costs nothing — that error is deliberately not billed, because detecting a missing text layer takes about 20 ms a page and charging for a "no" is how a feature earns refund requests.
{
"error": {
"code": "extraction_failed",
"message": "Every page is an image with no text layer. This looks like a scan, and this endpoint does not run OCR.",
"docs_url": "https://pdfcraft.dev/errors#extraction_failed"
}
}The honest recommendation
Try this first. If your documents are native text — and if they arrive by email, they almost certainly are — you get an answer in milliseconds that is the same every time and that you can verify against a bounding box.
For the scans, route to something that does OCR. A mixed pipeline is not a compromise, it is the correct design: the fast deterministic path for the 90% that does not need a model, and the expensive path only for the pages that do.
See it on a real document
The extraction playground takes a PDF and shows the JSON with every bounding box drawn over the page. No key needed for the first few.
Related reading
- How a PDF stores text, and why that makes extraction hard
- How PDF table extraction actually works
- How a table header is detected in a PDF
- PDF coordinates and bounding boxes, explained
- Deterministic PDF extraction versus an LLM
- What the confidence scores mean
- What it takes to run headless Chromium in production