PDFCraft
How it works

PDF coordinates and bounding boxes, explained

Every value returned carries the box it was read from. Getting that box to line up with what a browser would draw takes one flip most libraries do not do for you.

PDF space is upside down

A PDF page has its origin at the BOTTOM left, with y increasing upwards. Every screen coordinate system you have used — the DOM, canvas, every image format — has its origin at the top left with y increasing downwards.

So a raw PDF y coordinate is measured from the wrong end of the page. Draw a box at it directly over a rendered page image and it appears mirrored vertically: a value near the top of the document is drawn near the bottom of your overlay.

// The flip, which is all of it
y0 = viewportHeight - pdfY - glyphHeight;
y1 = viewportHeight - pdfY;

What you get back

Every bbox here is `[x0, y0, x1, y1]` in PDF points, origin TOP left, already flipped. A point is 1/72 inch, so an A4 page is 595 × 842 points.

That means a box can be laid over a rendered page with nothing but a scale factor — no flipping, no offset, no per-page arithmetic.

// Percentage positioning, which survives any render scale
const left   = (bbox[0] / pageWidth)  * 100;
const top    = (bbox[1] / pageHeight) * 100;
const width  = ((bbox[2] - bbox[0]) / pageWidth)  * 100;
const height = ((bbox[3] - bbox[1]) / pageHeight) * 100;

Why percentages rather than pixels

Position the overlay in percentages of the page, not in pixels of your canvas. A pixel-positioned box is correct only at the exact scale you computed it for, and any CSS that resizes the image — `max-width: 100%` on a narrow screen, say — moves the page under the boxes while leaving the boxes where they were.

That bug is easy to ship and hard to spot, because at the development width everything lines up.

Why this matters more than it sounds

A bounding box is the part of extraction that a language model cannot give you. It never sees coordinates, so it cannot tell you where on the page it read something — which means it cannot be checked.

With a box, a human reviewing an extraction can look at page 3 and see the number that was read. That is the difference between an answer you can audit and an answer you have to trust.

See it on a real document

The extraction playground takes a PDF and shows the JSON with every bounding box drawn over the page. No key needed for the first few.

Related reading