PDFCraft
How it works

Why this API does not do OCR, and how to tell if you need it

Adding OCR would destroy both things that make this useful: the latency and the determinism. Here is how to tell in one request whether your documents need it.

The two kinds of PDF

A PDF generated by software has a text layer — the actual characters, with positions. Reading it is a geometry problem and it takes milliseconds.

A PDF produced by a scanner or a phone camera is a picture of a document. There are no characters in the file at all, only pixels. Getting text out means recognising shapes, which is a fundamentally different operation with a fundamentally different cost.

Most business documents are the first kind. Invoices from accounting systems, statements from banks, reports from BI tools, anything emailed rather than posted — all native text.

What OCR would cost

Roughly 30 ms a page becomes roughly 2 seconds a page — a factor of sixty. Deterministic output becomes probabilistic: run the same file twice and the numbers can differ. And a clear "this has no text layer, send it elsewhere" becomes a confident, plausible, wrong answer.

That last one is the real objection. A parser that returns nothing is annoying. A parser that returns 8 where the document said 3 is a liability, and you find out about it downstream.

How to tell which you have, in one request

Send the file. If there is no text layer you get a 422 with the code `extraction_failed` and, in the body, exactly which pages were blank. It costs nothing — that error is deliberately not billed, because detecting a missing text layer takes about 20 ms a page and charging for a "no" is how a feature earns refund requests.

{
  "error": {
    "code": "extraction_failed",
    "message": "Every page is an image with no text layer. This looks like a scan, and this endpoint does not run OCR.",
    "docs_url": "https://pdfcraft.dev/errors#extraction_failed"
  }
}

The honest recommendation

Try this first. If your documents are native text — and if they arrive by email, they almost certainly are — you get an answer in milliseconds that is the same every time and that you can verify against a bounding box.

For the scans, route to something that does OCR. A mixed pipeline is not a compromise, it is the correct design: the fast deterministic path for the 90% that does not need a model, and the expensive path only for the pages that do.

See it on a real document

The extraction playground takes a PDF and shows the JSON with every bounding box drawn over the page. No key needed for the first few.

Related reading