ParseStream · PDF-to-HTML streaming
Does PDF-to-HTML stream like a real website?
ParseStream follows a PDF from session creation through the last readable block. It scores the HTTP surface, the shape of the stream, final-text proxies, routing decisions, and figure delivery—not only one finish time.
Snapshot updated August 12, 2026. Corpus: 19 documents across 8 classes.
Lane 0 · Website behavior · Passing
HTTP and streaming contract: 22 / 22 checks pass
A plain HTTP client can create a session, read useful no-JS content, resume the event stream, validate a finished document, issue HEAD, and observe liveness heartbeats.
Evidence boundary: One latest production lifecycle probe. The no-JS rendition contained about 99.2% of the finished page’s unique normalized words longer than two characters; this asymmetric recall check ignores order and frequency.
Lane 1 · Load trajectory · Measured
Streaming vitality · archived rubric: 3.4 s across-doc median
The E1 baseline exposes the real bottlenecks: first content misses the 2.5-second multi-page gate, large blocks can arrive in bursts, and figures can keep the page unfinished long after text lands.
Evidence boundary: The 3.4-second first-content median is from the original eight usable Lane 1 runs, one per document. In the later 15-run standing, documents passed 28.3% of an archived four-gate rubric that includes scale-invariant trajectory regret and received an F.
Grades · Routing · Directional
Document and figure routing: 19 docs · 175 labeled regions
The document policy matched 19/19 all-negative ground-truth calls. Across 175 recorded figure regions, MuPDF labels and pdf-inspector routing decisions agreed on embedded-raster versus fallback.
Evidence boundary: There is no true needs-vision positive. Figure F1 measures cross-engine agreement on image-XObject presence, not figure detection; in MuPDF-only documents, figures missed by MuPDF cannot enter the denominator.
Engine experiment · Provenance gap · Directional
Candidate engine A/B: Withheld · engine identity unproven
A local session produced two caller-labeled scorecards, but the harness did not pass an engine selector into the request. The results cannot be attributed to E1 or E2, so no performance comparison is published.
Evidence boundary: The artifacts omit session IDs, server-confirmed engine metadata, Worker version or commit, request configuration, and replay events. They remain public as failed-experiment evidence; a provenance-complete rerun is required before any A/B verdict.
Candidate-engine provenance table
- Engine-selecting request — published: No; required: Transmit the candidate selector to the service.
- Server-confirmed engine — published: No; required: Record session ID and terminal meta.engine.
- Worker build provenance — published: No; required: Record Worker version and source commit.
- Replayable run evidence — published: No; required: Publish request config and SSE event logs.
- Performance verdict — published: Withheld; required: Rerun only after every provenance cell is present.
Performance result withheld: the harness printed caller-provided engine labels but did not pass an engine selector into the request or record server-confirmed engine identity. Inspect the caller-labeled E1 scorecard and caller-labeled E2 scorecard.
GovTwin lexical text-proxy snapshot
These are normalized word 4-gram set overlap and Kendall tau-b over distinctive shared 6-gram positions. They do not measure semantic HTML, structure, alt text, accessibility, or complete reading order.
- Exceptions From Foreign Ownership (NRC rule) (Federal Register) — first content 3.5s; terminal event 9.4s; word 4-gram recall 90%; shared 6-gram order τ 1.000.
- Special Local Regulation; Genesee River (Federal Register) — first content 3.0s; terminal event 16.0s; word 4-gram recall 94%; shared 6-gram order τ 1.000.
- Medical Devices; Orthopedic Devices (Federal Register) — first content 3.4s; terminal event 35.6s; word 4-gram recall 93%; shared 6-gram order τ 1.000.
- Benefit and pension rates 2026–27 (GOV.UK) — first content 2.5s; terminal event 79.7s; word 4-gram recall 75%; shared 6-gram order τ 0.998.
Download the scored rows and read the exact scorer.
Limits and missing evidence
- Most performance rows are one live run per document. They are directional and should not be read as p50/p95 population estimates.
- The 19-document routing corpus contains scanned PDFs with publisher OCR, but no true no-text-layer document that should route to vision.
- Figure-routing F1 measures agreement between PDF engines on image-XObject presence for recorded regions. It does not score figure detection, and MuPDF-only fixtures cannot count figures that MuPDF missed.
- The image-layer result proves two delivery mechanisms, not a population percentage: almost all figures in the first experiment came from one raster-heavy PDF.
- The candidate-engine A/B cannot establish engine behavior: its headings use caller-provided labels, but the request did not transmit an engine selector and the artifacts contain no server-confirmed engine identity. No speed or quality result is published from it.
- The GovTwin table reports normalized word 4-gram overlap and the order of distinctive shared 6-gram anchors. Those lexical proxies do not establish semantic, structural, alt-text, or accessibility fidelity.
- Browser-level CLS, native image loading, Lighthouse, and axe form the planned Lane 2 and are not represented in the current standing.