Extract a utility bill PDF to JSON
Everything below is the response the live API returned for a utility bill, generated when this page was built. Not a description of what it would return — the output itself, including what it scored low on.
Who this is for
Energy-management, ESG-reporting and expense tools that need consumption figures, not just the total.
The job, in the words people search for: “read meter readings and charges off a utility bill”.
What makes this shape awkward
Two small tables — a tariff breakdown and a meter reading history — sit on one page with a charges summary. Telling three small tables apart from one big one is entirely a question of vertical gaps.
The measured result
| Pages read | 1 |
| Time | 11 ms (11 ms/page) |
| Labelled fields found unprompted | 4 |
| Tables | 2 (9 rows) |
| Input | utility-bill.pdf (29 KB — synthetic, generated from a spec in the repo; real documents of this type cannot be published) |
Fields, with a schema
Ask for the fields you want by the label printed on the page. Each answer carries the text exactly as printed (raw), the value coerced to the type you asked for, a confidence, and a bounding box you can draw on the page to check it.
curl -X POST https://api.pdfcraft.dev/v1/extract \
-H "Authorization: Bearer $PDFCRAFT_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"file": "<base64 of your utility bill>",
"schema": {
"account_number": "string",
"amount_due": "currency"
}
}'What came back
| Field | raw | value | confidence |
|---|---|---|---|
account_number | 4471-2209-88 | "4471-2209-88" | 60% |
amount_due | GBP 412.66 | 412.66 GBP | 70% |
Every label it found without being asked
Send no schema and you get all of these, keyed by the label as printed. Useful for discovering what a new supplier’s layout actually contains before you write a schema against it.
Account Number, Billing Period, Meter Serial, Amount Due
Tables
Description · Units · Rate · Amount · 3 rows × 4 columns · page 1
Table confidence 75%. Header confidence 100% — the first row was promoted out of the data.
| Description | Units | Rate | Amount |
|---|---|---|---|
| Electricity, day rate | 1,204 | 0.2841 | 342.06 |
| Electricity, night rate | 488 | 0.1402 | 68.42 |
| Standing charge | 90 | 0.5310 | 47.79 |
First 3 of 3 rows.
With options.rows_as_objects, the same row keyed by its header:
{
"Description": "Electricity, day rate",
"Units": "1,204",
"Rate": "0.2841",
"Amount": "342.06"
}Date · Reading · Type · 6 rows × 3 columns · page 1
Table confidence 90%. Header confidence 100% — the first row was promoted out of the data.
| Date | Reading | Type |
|---|---|---|
| 2026-01-01 | 41200 | Actual |
| 2026-01-15 | 41488 | Estimated |
| 2026-02-01 | 41776 | Actual |
First 3 of 6 rows.
With options.rows_as_objects, the same row keyed by its header:
{
"Date": "2026-01-01",
"Reading": "41200",
"Type": "Actual"
}What this does not do
- No OCR. This reads the PDF’s text layer. A scanned or photographed utility bill has no text layer and returns
422 extraction_failedwith the page numbers that were blank — deliberately, and quickly, so you can route it somewhere that does OCR rather than waiting on a guess. - No model. Run the same file twice and you get the same bytes back. That is the trade: it cannot infer a field that is not printed, and it cannot hallucinate one either.
- Ambiguity is reported, not resolved. A date like 03/04/2026 stays a string, and a bare
$returns a null currency, because guessing wrong on either is not something you can recover from downstream.
Try it on your own utility bill
The playground takes a file and shows the same JSON, with every bounding box drawn over the page. No key and no signup for the first few; a free key gives you 100 pages a month.