PDFCraft
How it works

How a PDF stores text, and why that makes extraction hard

A PDF does not contain words, lines, paragraphs or tables. It contains instructions for putting glyphs at coordinates. Everything else is inferred, by you or by whatever you point at the file.

What is actually in the file

A PDF page is a content stream: a sequence of operators that set a font, set a transformation matrix, and show a string. "Show this string at this position in this font at this size." That is the whole model.

There is no operator for "this is a paragraph". There is no operator for "this is a table cell". There is not even a reliable operator for "this is a space" — many producers advance the text position instead of emitting a space character, because moving the cursor is cheaper than drawing nothing.

So when you ask a library for the text of a page, it gives you an array of runs, each with a string and a box. Turning that into something you can use is the entire problem, and it is a geometry problem rather than a text problem.

// One text run, as it actually arrives
{
  text: "1,204",
  x0: 412.4, y0: 318.1, x1: 441.9, y1: 327.6,
  size: 9.5,
  font: "g_d0_f2"   // an id, not a name — but two runs in the same face share it
}

Why word spacing is guesswork

Runs frequently break mid-word. "Python, FastAPI, Node.js" with FastAPI in bold is three runs, and the boundaries fall exactly where the styling changes, not where the words do.

Join them with a space and you get "Python, FastAPI , Node.js" — close enough to look right, and wrong for anyone comparing the value against a known string. Join them without one and a genuine space between two runs disappears.

The only signal available is the gap. Below roughly 0.16 of the font size the runs are touching and it is a styling boundary; above it there is real whitespace. That threshold is measured, not chosen, and it is why extracted values here do not have stray spaces in them.

Why a column gap is different from a space

A space is about 0.25em. The gutter between two table columns is at least 1em, usually much more. That difference is the single idea the table reconstruction rests on: split a visual line wherever the gap exceeds 0.7em and you have separated columns from words, with no model and no training data.

Everything else follows. A table is a run of consecutive lines whose segments fall in the same vertical bands. A key/value block is a two-column table whose first column always ends in a colon. A bullet list is a two-column table whose first column is always a dash.

What this means for you

If your PDF was generated by software — an invoice from an accounting system, a statement from a bank, a report from a BI tool — the text layer is there and clean, and geometry alone gets you a long way, fast and deterministically.

If it was scanned or photographed, there is no text layer at all, and no amount of geometry helps. That is OCR’s job, and it is a different kind of tool with a different cost and a different error profile.

See it on a real document

The extraction playground takes a PDF and shows the JSON with every bounding box drawn over the page. No key needed for the first few.

Related reading