Bill of lading OCR
Nobody had published logistics-document numbers, so we built the set: the same 25 bills of lading and the same 21-field schema through every parser, split by digital, scanned, and fax-quality condition.
24 systems · 25 BoL PDFs · 21 fields · 3 conditions · last run August 18, 2026 · benchmarks/bol-bench · Bill of lading benchmark writeup.
Winning result
Quality — 96.9% (okraPDF /v1)
Value — not scored on this board.
Speed — 0.02s (PyMuPDF)
Leaderboard
| # | System | Field accuracy | Digital | Scanned | Fax quality | Median parse |
| 1 | okraPDF /v1 Vendor API | 96.9% | 100.0% | 100.0% | 84.2% | 29.39s |
|---|
| 2 | Mistral OCR Vendor API | 96.6% | 98.5% | 100.0% | 86.8% | 1.84s |
|---|
| 3 | google/gemini-3.1-flash-lite VLM | 94.9% | 99.0% | 100.0% | 77.2% | 2.90s |
|---|
| 4 | qwen/qwen3-vl-235b-a22b-instruct VLM | 94.9% | 100.0% | 100.0% | 74.1% | 25.22s |
|---|
| 5 | anthropic/claude-sonnet-5 VLM | 94.5% | 100.0% | 100.0% | 72.4% | 10.49s |
|---|
| 6 | qwen/qwen3-vl-8b-instruct VLM | 93.5% | 98.5% | 100.0% | 71.1% | 6.70s |
|---|
| 7 | LlamaParse Vendor API | 92.1% | 98.5% | 100.0% | 64.5% | 22.09s |
|---|
| 8 | Reducto Vendor API | 90.8% | 98.5% | 100.0% | 57.9% | 65.89s |
|---|
| 9 | z-ai/glm-4.6v VLM | 89.9% | 99.0% | 87.7% | 68.9% | 40.14s |
|---|
| 10 | google/gemini-3.7-flash VLM | 89.8% | 100.0% | 100.0% | 48.7% | 8.56s |
|---|
| 11 | mistralai/mistral-small-2603 VLM | 89.3% | 98.0% | 100.0% | 51.3% | 5.95s |
|---|
| 12 | meta-llama/llama-4-maverick VLM | 89.1% | 96.5% | 100.0% | 54.4% | 6.37s |
|---|
| 13 | Chandra (Datalab) Vendor API | 88.9% | 98.0% | 100.0% | 49.6% | 7.99s |
|---|
| 14 | moonshotai/kimi-k2.5 VLM | 88.8% | 92.0% | 100.0% | 64.9% | 22.39s |
|---|
| 15 | openai/gpt-5-mini VLM | 88.0% | 100.0% | 100.0% | 39.5% | 22.43s |
|---|
| 16 | google/gemini-2.5-flash VLM | 87.6% | 92.5% | 87.7% | 74.6% | 4.55s |
|---|
| 17 | anthropic/claude-haiku-4.5 VLM | 86.4% | 93.3% | 99.1% | 50.4% | 7.68s |
|---|
| 18 | openai/gpt-5.4-mini VLM | 85.1% | 99.5% | 100.0% | 26.3% | 3.26s |
|---|
| 19 | amazon/nova-2-lite-v1 VLM | 84.1% | 94.0% | 97.5% | 39.5% | 3.90s |
|---|
| 20 | openai/gpt-4o-mini VLM | 83.3% | 90.7% | 94.3% | 48.7% | 7.13s |
|---|
| 21 | Tesseract 5 OCR | 79.1% | 96.5% | 90.6% | 17.1% | 1.46s |
|---|
| 22 | pdftotext (poppler) Text layer | 52.4% | 100.0% | 0.0% | 0.0% | 0.03s |
|---|
| 23 | PyMuPDF Text layer | 52.4% | 100.0% | 0.0% | 0.0% | 0.02s |
|---|
| 24 | pdfplumber Text layer | 52.1% | 99.5% | 0.0% | 0.0% | 0.05s |
- We built this corpus and we score on it, so read the okraPDF row with that in mind. Mistral OCR wins fax-quality scans and is 16x faster at the median.
- Every pipeline parsed all 25 documents without error; the text-layer parsers score 0 on scans because there is no text layer to read.