HTML to PDF in Java
One POST, a PDF back. No headless browser to install, no Chromium to keep patched, no fonts to install on the box — the render happens on a warm browser that is already running.
Render HTML to a PDF
import java.net.URI;
import java.net.http.*;
import java.nio.file.Path;
import java.time.Duration;
var body = """
{"html":"<h1>Invoice 1042</h1>","options":{"printBackground":true}}
""";
var request = HttpRequest.newBuilder(URI.create("https://api.pdfcraft.dev/v1/render"))
.header("Authorization", "Bearer " + System.getenv("PDFCRAFT_API_KEY"))
.header("Content-Type", "application/json")
.timeout(Duration.ofSeconds(60))
.POST(HttpRequest.BodyPublishers.ofString(body))
.build();
var client = HttpClient.newHttpClient();
var response = client.send(request, HttpResponse.BodyHandlers.ofFile(Path.of("invoice.pdf")));
if (response.statusCode() != 200) {
throw new IllegalStateException("render failed: " + response.statusCode());
}Read a PDF back as JSON
The same key works for extraction. Send a PDF, name the fields you want by the label printed on the page, and get them back with the raw text, a coerced value, a confidence and a bounding box.
var encoded = Base64.getEncoder().encodeToString(Files.readAllBytes(Path.of("statement.pdf")));
var payload = new ObjectMapper().writeValueAsString(Map.of(
"file", encoded,
"schema", Map.of("account_number", "string", "closing_balance", "currency")
));
var request = HttpRequest.newBuilder(URI.create("https://api.pdfcraft.dev/v1/extract"))
.header("Authorization", "Bearer " + System.getenv("PDFCRAFT_API_KEY"))
.header("Content-Type", "application/json")
.POST(HttpRequest.BodyPublishers.ofString(payload))
.build();
var response = client.send(request, HttpResponse.BodyHandlers.ofString());
var json = new ObjectMapper().readTree(response.body());
System.out.println(json.at("/fields/closing_balance/value"));What bites in Java
- `BodyHandlers.ofFile` streams the response straight to disk without ever holding the PDF in the heap. For a large render that is the difference between a few kilobytes of buffer and a heap spike.
- `Base64.getEncoder()`, not `getMimeEncoder()` — the MIME encoder line-wraps.
- HttpClient is thread-safe and expensive to build. Make one and reuse it; a client per request defeats connection pooling and is the usual cause of "it is slower from Java".
Errors are one shape
Every failure is {"error":{"code","message","docs_url"}} with a stable code, so you can switch on error.code rather than parsing prose. The two worth handling explicitly are rate_limited — honour Retry-After — and render_failed, which means your HTML broke rather than ours did.
The same thing in another language
- HTML to PDF in Node.js
- HTML to PDF in Python
- HTML to PDF in Go
- HTML to PDF in PHP
- HTML to PDF in Ruby
- HTML to PDF in C#
- HTML to PDF in Rust
- HTML to PDF in curl
- HTML to PDF in Deno
- HTML to PDF in Bun
- HTML to PDF in TypeScript
Or try it with no code at all in the playground. A free key is 100 renders a month.