PDFCraft
Extraction guide

Extract a insurance claim form PDF to JSON

Everything below is the response the live API returned for a insurance claim form, generated when this page was built. Not a description of what it would return — the output itself, including what it scored low on.

Who this is for

Insurtech and TPA platforms where the intake channel is still a filled-in PDF form.

The job, in the words people search for: “read completed claim forms into a claims system”.

What makes this shape awkward

Form fields whose values are typed into boxes below the label rather than beside it. The value is on the next visual line, which the label-and-value rule has to reach down for without swallowing the following section.

The measured result

Pages read1
Time14 ms (14 ms/page)
Labelled fields found unprompted6
Tables1 (7 rows)
Inputinsurance-claim.pdf (29 KB — synthetic, generated from a spec in the repo; real documents of this type cannot be published)

Fields, with a schema

Ask for the fields you want by the label printed on the page. Each answer carries the text exactly as printed (raw), the value coerced to the type you asked for, a confidence, and a bounding box you can draw on the page to check it.

curl -X POST https://api.pdfcraft.dev/v1/extract \
  -H "Authorization: Bearer $PDFCRAFT_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "file": "<base64 of your insurance claim form>",
    "schema": {
      "claim_number": "string",
      "policy_number": "string",
      "date_of_loss": "date",
      "estimated_value": "currency"
    }
  }'

What came back

Fieldrawvalueconfidence
claim_numberCLM-2026-338914"CLM-2026-338914"60%
policy_numberPOL-77-441290"POL-77-441290"60%
date_of_loss2026-02-19"2026-02-19"70%
estimated_valueEUR 27,400.0027400 EUR70%

Every label it found without being asked

Send no schema and you get all of these, keyed by the label as printed. Useful for discovering what a new supplier’s layout actually contains before you write a schema against it.

Claim Number, Policy Number, Date of Loss, Date Reported, Claim Type, Estimated Value

Tables

Item · Description · Purchase date · Claimed · 7 rows × 4 columns · page 1

Table confidence 90%. Header confidence 100% — the first row was promoted out of the data.

ItemDescriptionPurchase dateClaimed
1Laptop2023-01-101,347.00
2Monitor, 32in2024-02-11843.00
3Desk2025-03-122,265.00

First 3 of 7 rows.

With options.rows_as_objects, the same row keyed by its header:

{
  "Item": "1",
  "Description": "Laptop",
  "Purchase date": "2023-01-10",
  "Claimed": "1,347.00"
}

What this does not do

Try it on your own insurance claim form

The playground takes a file and shows the same JSON, with every bounding box drawn over the page. No key and no signup for the first few; a free key gives you 100 pages a month.

Related