PDFCraft
Scanned documents

Are scanned PDFs accessible?

No. A scanned PDF contains an image of a page, not text. A screen reader opens it and finds nothing to announce; the content cannot be searched, copied, translated or reflowed on a phone. It is the most complete form of inaccessibility a document can have, and the most expensive to undo.

What a scanned PDF actually is

When a page goes through a scanner or a photocopier with a PDF output, what comes out is a photograph. The letters you can see are pixels arranged in the shape of letters. Nothing in the file says “this says Patient Financial Assistance Application” — that meaning exists only in the eye of someone looking at it.

This is different from, and worse than, an untagged PDF. An untagged text PDF at least contains the words, so a screen reader can read them out in a flat, structureless stream. A scan contains no words at all.

How to tell, quickly

Open it and try to select a line of text with your cursor. If you cannot, or if you can only select the whole page as one rectangle, it is a scan. Searching for a word you can clearly see on screen is the same test.

A middle case exists and is worth knowing about: a scan that has had OCR applied has an invisible text layer behind the image. It will pass the selection test. But OCR output is frequently wrong in ways nobody has checked, and it almost never carries structure — no headings, no table relationships, no reading order. A screen reader will read something, and that something may not be what is on the page.

Detecting them across thousands of documents

You cannot open five thousand PDFs. Fortunately this is the one accessibility question with a completely reliable automated answer: either the content stream contains extractable glyphs or it does not.

PDFCraft answers it as a side effect of its other job. The same extraction engine that turns PDFs into structured JSON refuses documents with no text layer, because there is nothing to extract — and that refusal is precisely the accessibility finding. It is reported as a blocker when the whole document is a scan.

Partly scanned documents get their own treatment, and they matter more than they look. A report with scanned appendices reads normally and then goes silent partway through, with nothing to announce that anything is missing — harder for a reader to diagnose than a document that fails outright. Those are reported with the specific page numbers.

Every document in the 15-document federal corpus we scanned on September 27, 2026 has a text layer, which is what you would expect from born-digital federal publishing. Scans cluster somewhere else entirely: older municipal records, board minutes, signed forms, anything that passed through a fax or a photocopier on its way to the website.

Why they dominate a remediation budget

Remediating a tagged document is correction: the structure exists and something about it is wrong. Remediating a scan is construction. The sequence is roughly:

  1. Recognise the text (OCR), which is fast and cheap.
  2. Proofread the recognition, which is neither — OCR confuses similar glyphs, mangles tables, and does unpredictable things to handwriting, stamps and form fields.
  3. Build a tag structure from nothing: headings, lists, tables, figures.
  4. Decide the reading order. For a multi-column or form layout this is a judgment a person has to make, not a setting.
  5. Write alternative text for anything that stays an image.

That is why a scanned page sits at the top of the $5–$25 per page range, and why our cost model counts it at 2.5× a plain page. In an estate with a meaningful proportion of scans, they routinely account for more of the estimate than every other document combined — you can watch that happen in the cost calculator by moving the scan percentage.

The cheapest fix is usually not remediation

For scanned documents more than any other category, it is worth asking whether the document needs to exist in that form at all before paying to rebuild it.

Triage exists to ask these questions across a whole estate before anyone is paid per page. That is the entire argument for doing it first.

What we do not do

We do not perform OCR and we do not remediate documents. PDFCraft finds your scanned documents, tells you how many there are, which ones people actually reach, and what the remediation range looks like. Fixing them is a service business with established vendors, and the report is what lets you scope that engagement instead of opening it.

Related