How PDF extraction and rendering actually work
The mechanism, written down. Most of this came out of building the extractor and fixing what measurement found, which is why it is specific rather than general.
- How a PDF stores text, and why that makes extraction hard — A PDF does not contain words, lines, paragraphs or tables. It contains instructions for putting glyphs at coordinates. E…
- How PDF table extraction actually works — No machine learning, no ruling-line detection, no training set. Tables are found by measuring horizontal gaps and vertic…
- How a table header is detected in a PDF — A header row looks exactly like a data row. Nothing in the file marks it. Here is the evidence that is actually availabl…
- Why this API does not do OCR, and how to tell if you need it — Adding OCR would destroy both things that make this useful: the latency and the determinism. Here is how to tell in one …
- PDF coordinates and bounding boxes, explained — Every value returned carries the box it was read from. Getting that box to line up with what a browser would draw takes …
- Deterministic PDF extraction versus an LLM — Not a rivalry. They fail differently, cost differently by two orders of magnitude, and the sensible architecture uses bo…
- What the confidence scores mean — Every field and every table comes back with a number between 0 and 1. Here is exactly what goes into each one, so you ca…
- What it takes to run headless Chromium in production — Puppeteer works in ten lines. The gap between that and something a service can depend on is where the time goes — and it…