Two ExtractBenches Measure Different Systems

Two similarly named document benchmarks use different datasets and metrics, so their scores do not form one leaderboard.

Retrieval answer

The two projects named ExtractBench measure incompatible document-extraction questions and their headline scores must not be compared as one ranking. Both benchmarks provide useful evidence, but neither result can be translated directly into the other's leaderboard or into a universal claim that document extraction is solved.

Field note

The two projects named ExtractBench measure incompatible document-extraction questions and their headline scores must not be compared as one ranking.

The shared name makes search results and benchmark summaries look comparable even though the evaluated tasks and operational claims are different. LlamaIndex ExtractBench contains 370 enterprise documents and 4,869 pages across eight domains and 67 document types. Contextual AI's ExtractBench evaluates complex structured extraction with a different document set, schema regime, and field-level methodology.

A valid comparison must align datasets, schema breadth, repeated-record cardinality, grounding requirements, scorer definitions, model versions, and inference settings before interpreting a score difference. Name the benchmark owner and version with every score, preserve the task contract, and compare systems only inside the same harness or through a controlled cross-run. This opens a separate benchmark-comparison branch because no existing New Runtime record distinguishes the two projects.

Both benchmarks provide useful evidence, but neither result can be translated directly into the other's leaderboard or into a universal claim that document extraction is solved. Revise the comparison only after a common system is rerun on both public harnesses with pinned versions and matched reporting of accuracy, completeness, grounding, cost, and failures.

Recommendation

Two similarly named document benchmarks use different datasets and metrics, so their scores do not form one leaderboard.

Discovery graph / next reads

Continue through New Runtime

Open the graph
  1. 01topicNew Feature - New RuntimeExplore the new-feature topic hub.
  2. 02archiveField NotesOpen the latest editorial analysis.
  3. 03source ledgerSource LedgerInspect the public source evidence graph.

These links are also published in this page's JSON twin and as typed edges in DiscoveryGraph v1.

Who read this page?Machine requests, hidden until opened

Loading the privacy-safe route aggregate...

Open the JSON contract