_

Benchmarks for the hardest parts of document processing

Published document processing benchmarks for parsing, structured extraction, and splitting. Each report states its corpus, metric, comparison scope, methodology, and public data or code.

_

Benchmarks

Public benchmarks for document parsing, extraction, and splitting, with published methodology and source data. Each card separates key results, benchmark details, and sources.

Parsing

RealDoc-Bench

Measures whether parsers deliver accurate layouts, preserve reading order, and enable agents to correctly answer objective questions against real-world documents.

Q&A accuracy
95.7%
Extend Parse 2.0
Layout F1
0.847
Adjusted F1
Benchmark details for RealDoc-Bench
Benchmark details for RealDoc-Bench
FieldDetail
Test corpus1,500 layout samples; 1,359 document Q&A prompts across 581 documents in four regulated industries
MetricAdjusted F1 for layout; field-level document Q&A accuracy
ComparisonSame-run comparison with LlamaParse, Reducto, Azure Document Intelligence, and AWS Textract
PublishedMay 26, 2026

Sources

Read RealDoc-Bench

Extraction

LongArray-Extract

Tests whether extraction systems preserve cardinality and return complete, schema-faithful arrays when the output grows from a dozen of rows to thousands.

Mean accuracy
99.2%
Extend across 45 PDFs
Run completion
100%
45 of 45 PDFs completed
Benchmark details for LongArray-Extract
Benchmark details for LongArray-Extract
FieldDetail
Test corpus45 financial, clinical, and legal PDFs with repeated arrays of 27 to about 2,200 records
MetricMean per-document extraction accuracy, with failed and timed-out runs scored as zero
Comparison
  • Document-AI platforms: Reducto, Pulse, LlamaParse
  • Raw-model providers: Anthropic, Google, OpenAI
PublishedJune 2, 2026

Sources

Read LongArray-Extract

Splitting

PoliTax Split

Evaluates document splitting on long, compound tax filings where frontier models miss subtle boundaries across hundreds of pages.

Best harness F1
72.48%
Claude Opus 4.6
Recall lift
17-44
Points across models
Benchmark details for PoliTax Split
Benchmark details for PoliTax Split
FieldDetail
Test corpusThe 30 largest compound documents in the public PoliTax corpus
MetricBoundary-detection F1
ComparisonSame-model comparison between the Extend splitting harness and direct frontier-model use
PublishedMarch 30, 2026

Sources

Read PoliTax Split
cta-background

( fig.11 )

Turn your documents into high quality data