Field note
The two projects named ExtractBench measure incompatible document-extraction questions and their headline scores must not be compared as one ranking.
The shared name makes search results and benchmark summaries look comparable even though the evaluated tasks and operational claims are different. LlamaIndex ExtractBench contains 370 enterprise documents and 4,869 pages across eight domains and 67 document types. Contextual AI's ExtractBench evaluates complex structured extraction with a different document set, schema regime, and field-level methodology.
A valid comparison must align datasets, schema breadth, repeated-record cardinality, grounding requirements, scorer definitions, model versions, and inference settings before interpreting a score difference. Name the benchmark owner and version with every score, preserve the task contract, and compare systems only inside the same harness or through a controlled cross-run. This opens a separate benchmark-comparison branch because no existing New Runtime record distinguishes the two projects.
Both benchmarks provide useful evidence, but neither result can be translated directly into the other's leaderboard or into a universal claim that document extraction is solved. Revise the comparison only after a common system is rerun on both public harnesses with pinned versions and matched reporting of accuracy, completeness, grounding, cost, and failures.