PDFCraft
PDF accessibility triage

How many of your PDFs are actually accessible?

Point it at a domain. It finds the PDFs, checks each one, and returns a prioritised report: what fails, how severely, what remediation would cost, and what to fix first. The first 25 documents are free and need no account.

The problem is not finding failures. It is knowing which ones matter.

Free single-file checkers already exist, and they will tell you that a document fails. What they will not tell you is which of your five thousand documents to fix on Monday. A missing document title and an unlabelled form field are both conformance failures. One costs a reader a few seconds of confusion; the other means the form cannot be completed at all. A report that lists them in the same voice is a list, not a plan — and a list does not get a budget approved.

So the output here is ranked by consequence and priced. Severity is blocker (the document cannot be used), major (usable, but a specific task will fail), or minor (a real defect measured in seconds). Priority is severity weighted by reach, because a blocker on a document linked from your homepage matters more than the same blocker on a 2014 board minute.

A scan of 15 federal documents, run on September 27, 2026

Published documents from the IRS, Congress, the Census Bureau, the CDC, the Federal Reserve and the Department of Justice — 492 pages in total, all public, all produced by well-resourced organisations with accessibility obligations of their own.

15 of 15 failed PDF/UA conformance. 3 have no tag structure at all, and 10 contain at least one blocker. Among the documents that fail is the published text of the ADA Title II web accessibility regulation itself.

That last fact is the most useful one on this page, and it is not here as a gotcha. It sets a realistic expectation: if you scan your own portfolio and most of it fails, you have a normal document estate, not a negligent one. Nobody’s portfolio is clean, which is exactly why triage has to come before remediation.

DocumentPublisherPagesTaggedBlockersMajorMinor
Economic Well-Being of U.S. HouseholdsFederal Reserve82no266
H.R. 815 (enrolled)U.S. Congress110no455
28 CFR §35.105 — the ADA Title II web accessibility ruleDepartment of Justice1no465
Form 1040, U.S. Individual Income Tax ReturnIRS2yes251
Form 1099-MISC, Miscellaneous InformationIRS6yes250
Form 4506-T, Request for Transcript of Tax ReturnIRS2yes031
Form 8822, Change of AddressIRS2yes345
Form 941, Employer's Quarterly Federal Tax ReturnIRS3yes251
Form SS-4, Application for Employer Identification NumberIRS2yes220
Form W-4, Employee's Withholding CertificateIRS5yes341
Form W-9, Request for Taxpayer Identification NumberIRS6yes231
Form 1040 InstructionsIRS126yes086
National Vital Statistics Report 72-11CDC / NCHS19yes036
Publication 15, Employer's Tax GuideIRS59yes044
P60-280, Income in the United StatesU.S. Census Bureau67yes0108

Scanned with pdfa11y 0.0.11 for machine conformance, plus PDFCraft’s own geometric checks. Machine conformance checks from an open-source PDF/UA validator, plus geometric checks run against the same extraction engine the rest of this API uses. Every document is public and published by the body named beside it.

Two layers, because half the problem is not machine-checkable

The PDF Association’s Matterhorn Protocol splits its checkpoints into machine-verifiable and human-judgment categories, and the split is the whole story. A document can pass every automated conformance check and still be unusable with a screen reader.

The first layer is an open-source PDF/UA validator, which answers the machine-verifiable questions: is there a structure tree, do images carry alternative text, do table headers declare a scope, do form fields have labels. Across this corpus it found 39 distinct failing checks.

The second layer is ours, and it exists because a validator can only check that something is present, not that it is right. It found issues on 8 of the 15 documents above:

What we checkWhy a conformance validator cannot
No text layer — the page is an image of textIt validates structure, not whether the page is a picture. Our extraction engine already rejects scans, and that rejection is the finding — it is also the most expensive category to remediate.
Reading order does not match the pageA validator can prove that an order exists, not that it is sensible. We recover the visual reading order from the page geometry and compare it with the order stored in the tags, so a page whose tags run bottom-to-top is caught even though it looks perfect on screen.
Tagged as a table but not laid out as oneA layout table with header cells passes. We test whether the cells actually line up on a grid, under left, right or centre alignment.
A real table carrying no table tagGrid-aligned content with no table tag reads as a stream of loose numbers with every row and column heading stripped off.
Alt text that does not describe anythingPresence of alt text is machine-checkable; quality is not. We flag values that are filenames, placeholder words, too short to describe anything, or pasted onto every figure in the document.

Reading order is the one worth dwelling on, because it is invisible to everyone who is not using a screen reader. A sighted reviewer opens the page, sees the columns in the right places, and signs it off. The file can still say “read the right column first”, and every automated tool on the market will call it conformant.

What it costs

Remediation is quoted publicly by established vendors at roughly $5 to $25 per page, and we report an estimate as a range rather than a single number because the spread is genuine — a tagged document needing a title is not the same job as a scan needing rebuilding. The cost calculator puts a range around your own estate.

We do not remediate documents. Fixing PDFs is a service business with ten established vendors and real expertise; we are the step before you call one. Triage is what turns an open-ended engagement into a scoped one.

Pricing

PlanDocuments per scanPrice
Free scan25Free, no card
One-off report1,000$49 once
Monitor10,000$149/month
Scale50,000$399/month

The free tier is a genuine scan of 25 documents with full findings, not a preview with the answers removed. It is the top of the funnel and it is no use to anyone crippled.

The deadlines

Both sets moved during 2026, in two separate interim final rules six months apart, and a great deal of published guidance is still wrong about at least one of them. The current dates:

WhoDeadlineRule
Public entities serving a population of 50,000 or moreApril 26, 2027ADA Title II
Recipients of HHS funding with 15 or more employees — hospitals, health systems, most clinicsMay 11, 2027HHS Section 504
Public entities serving under 50,000, and special districtsApril 26, 2028ADA Title II
Recipients of HHS funding with fewer than 15 employeesMay 10, 2028HHS Section 504

Verified September 27, 2026 against the Federal Register. Full detail, with what each date used to be and the document number it changed in, on the deadlines page.

Read next