PDF parsing benchmarks
Independent, reproducible measurements of document parsers, OCR systems, and vision models. Every score names its corpus, its harness, and the configuration behind it.
6 benchmarks · 48 systems scored · last run September 9, 2026.
Parsing quality
Page-level parse quality on public harnesses: tables, charts, faithfulness, formatting, and visual grounding.
| Benchmark | Quality | Value | Speed |
|
ParseBench — hosted parsers
The published ParseBench leaderboard: twelve hosted parsing pipelines scored on tables, charts, content faithfulness, semantic formatting, and visual grounding.
12 systems · ~2,000 human-verified enterprise pages · last run August 6, 2026
|
84.88 · LlamaParse Agentic |
0.28¢ · Databricks AI Parse |
— |
|
ParseBench, run locally
Open parsers scored on the official ParseBench harness on our own hardware: five dimensions, one seeded page subset, bootstrap intervals, measured throughput and GPU memory. Rule-based scoring, judge off.
2 systems · 60 seeded pages per dimension · AMD Radeon RX 7900 XT 20 GB · last run September 9, 2026
|
52.4 · PyMuPDF4LLM + RapidOCR |
— |
0.9s · PyMuPDF4LLM + RapidOCR |
|
Agentic parsing on ParseBench
Can a text-only agent driving crop and ask-a-VLM tools beat the single-shot vision model it calls? Scored on the official harness against that VLM and a Gemini Flash anchor.
5 systems · 6 pages · 3 chart + 3 table docs · last run August 16, 2026
|
50% · Gemini 3 Flash (anchor) |
$0 actual · DeepSeek agent 1800s |
— |
Extraction & OCR
Field-level and line-level reads, including the logistics documents and handwriting lines no public board covered.
| Benchmark | Quality | Value | Speed |
|
Bill of lading OCR
Nobody had published logistics-document numbers, so we built the set: the same 25 bills of lading and the same 21-field schema through every parser, split by digital, scanned, and fax-quality condition.
24 systems · 25 BoL PDFs · 21 fields · 3 conditions · last run August 18, 2026
|
96.9% · okraPDF /v1 |
— |
0.02s · PyMuPDF |
|
Handwriting recognition — IAM test split
Four systems over the same handwritten lines from the canonical IAM-line test split, scored with the same character and word error rates.
4 systems · 120 lines · IAM-line test split · last run June 22, 2026
|
1.29% · okraPDF /v1 |
— |
— |
Delivery & streaming
Whether a live PDF-to-HTML twin behaves like a website — contract, first content, and reading order — not only final text.
| Benchmark | Quality | Value | Speed |
|
PDF-to-HTML streaming (ParseStream)
One pipeline followed from session creation to the last readable block: HTTP contract, load trajectory, routing, figure delivery, and text overlap against government HTML twins.
1 system · 19 docs · 8 classes · 15 live runs · last run August 12, 2026
|
22 / 22 · HTTP contract |
τ 1.00 · vs the publisher’s own HTML |
3.4 s · first content · E1 single-call |
Methods
Every board above names its corpus, its harness, and the date it last ran. Two caller-labeled candidate-engine scorecards are published as provenance-gap evidence, not as an A/B result. Or run PDF to HTML yourself.