HTML to PDF in Python
One POST, a PDF back. No headless browser to install, no Chromium to keep patched, no fonts to install on the box — the render happens on a warm browser that is already running.
With the official client
There is a published Python client, so the shortest version is two lines. It has no dependencies of its own, retries 429s and 5xxs for you, and turns the API’s error envelope into a typed error you can branch on. The raw HTTP below still works and is still supported — plenty of people would rather send a POST than take a dependency.
pip install pdfcraftfrom pdfcraft import PDFCraft
client = PDFCraft(os.environ["PDFCRAFT_API_KEY"])
pdf = client.render(html="<h1>Invoice 1042</h1>")
open("invoice.pdf", "wb").write(pdf)
data = client.extract_pdf(open("statement.pdf", "rb").read())
print(data["tables"][0]["header"])On PyPI.
Render HTML to a PDF with plain HTTP
import os, requests
response = requests.post(
"https://api.pdfcraft.dev/v1/render",
headers={"Authorization": f"Bearer {os.environ['PDFCRAFT_API_KEY']}"},
json={
"html": "<h1>Invoice 1042</h1>",
"options": {"printBackground": True, "margin": {"top": "20mm", "bottom": "20mm"}},
},
timeout=60,
)
response.raise_for_status()
with open("invoice.pdf", "wb") as handle:
handle.write(response.content)Read a PDF back as JSON
The same key works for extraction. Send a PDF, name the fields you want by the label printed on the page, and get them back with the raw text, a coerced value, a confidence and a bounding box.
import base64, os, requests
with open("statement.pdf", "rb") as handle:
encoded = base64.b64encode(handle.read()).decode()
response = requests.post(
"https://api.pdfcraft.dev/v1/extract",
headers={"Authorization": f"Bearer {os.environ['PDFCRAFT_API_KEY']}"},
json={
"file": encoded,
"schema": {"account_number": "string", "closing_balance": "currency"},
"options": {"rows_as_objects": True},
},
timeout=60,
)
body = response.json()
print(body["fields"]["closing_balance"]["value"])
for row in body["tables"][0]["rows_as_objects"]:
print(row)What bites in Python
- `response.content`, not `response.text`. `.text` decodes the bytes as a string and the PDF is destroyed — this is the single most common way a Python integration produces a file that will not open.
- Pass `timeout=`. requests has no default timeout at all, so a hung connection hangs your worker indefinitely rather than raising.
- For async use httpx with the same shape; the API has no streaming requirement, so the sync client is fine for most jobs.
Errors are one shape
Every failure is {"error":{"code","message","docs_url"}} with a stable code, so you can switch on error.code rather than parsing prose. The two worth handling explicitly are rate_limited — honour Retry-After — and render_failed, which means your HTML broke rather than ours did.
The same thing in another language
- HTML to PDF in Node.js
- HTML to PDF in Go
- HTML to PDF in PHP
- HTML to PDF in Ruby
- HTML to PDF in Java
- HTML to PDF in C#
- HTML to PDF in Rust
- HTML to PDF in curl
- HTML to PDF in Deno
- HTML to PDF in Bun
- HTML to PDF in TypeScript
Or try it with no code at all in the playground. A free key is 100 renders a month.