PDFCraft
Extraction guide

Extract a work order PDF to JSON

Everything below is the response the live API returned for a work order, generated when this page was built. Not a description of what it would return — the output itself, including what it scored low on.

Who this is for

Maintenance and field-service platforms replacing paper job sheets that come back scanned or exported.

The job, in the words people search for: “read maintenance work orders into a CMMS”.

What makes this shape awkward

Two small tables — parts used and labour hours — with different column counts, under separate headings, on one page. Each needs its own column structure, and measuring one against the other loses a header.

The measured result

Pages read1
Time10 ms (10 ms/page)
Labelled fields found unprompted6
Tables2 (7 rows)
Inputwork-order.pdf (29 KB — synthetic, generated from a spec in the repo; real documents of this type cannot be published)

Fields, with a schema

Ask for the fields you want by the label printed on the page. Each answer carries the text exactly as printed (raw), the value coerced to the type you asked for, a confidence, and a bounding box you can draw on the page to check it.

curl -X POST https://api.pdfcraft.dev/v1/extract \
  -H "Authorization: Bearer $PDFCRAFT_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "file": "<base64 of your work order>",
    "schema": {
      "work_order": "string",
      "raised": "date",
      "completed": "date"
    }
  }'

What came back

Fieldrawvalueconfidence
work_orderWO-2026-8841"WO-2026-8841"60%
raised2026-03-08"2026-03-08"70%
completed2026-03-11"2026-03-11"70%

Every label it found without being asked

Send no schema and you get all of these, keyed by the label as printed. Useful for discovering what a new supplier’s layout actually contains before you write a schema against it.

Work Order, Asset, Priority, Raised, Completed, Technician

Tables

Part number · Description · Qty · 4 rows × 3 columns · page 1

Table confidence 75%. Header confidence 100% — the first row was promoted out of the data.

Part numberDescriptionQty
PN-4412Air filter element2
PN-7719Oil seal, 40mm1
PN-2204Drive belt1

First 3 of 4 rows.

With options.rows_as_objects, the same row keyed by its header:

{
  "Part number": "PN-4412",
  "Description": "Air filter element",
  "Qty": "2"
}

Date · Technician · Hours · Rate · 3 rows × 4 columns · page 1

Table confidence 75%. Header confidence 100% — the first row was promoted out of the data.

DateTechnicianHoursRate
2026-03-08S. Idris3.568.00
2026-03-09S. Idris6.068.00
2026-03-11A. Pike2.074.00

First 3 of 3 rows.

With options.rows_as_objects, the same row keyed by its header:

{
  "Date": "2026-03-08",
  "Technician": "S. Idris",
  "Hours": "3.5",
  "Rate": "68.00"
}

What this does not do

Try it on your own work order

The playground takes a file and shows the same JSON, with every bounding box drawn over the page. No key and no signup for the first few; a free key gives you 100 pages a month.

Related