---
type: "post"
slug: "two-extractbenches-different-systems"
title: "Two ExtractBenches Measure Different Systems"
description: "Two similarly named document benchmarks use different datasets and metrics, so their scores do not form one leaderboard."
retrieval_nugget: "The two projects named ExtractBench measure incompatible document-extraction questions and their headline scores must not be compared as one ranking. Both benchmarks provide useful evidence, but neither result can be translated directly into the other's leaderboard or into a universal claim that document extraction is solved."
published_at: "2026-08-27"
updated_at: "2026-08-27"
record_date: "2026-08-27"
date_kind: "published_at"
topics: ["new-feature"]
entities: []
source_urls: ["https://extractbench.ai/","https://arxiv.org/abs/2607.29677","https://arxiv.org/abs/2602.12247","https://github.com/ContextualAI/extract-bench"]
source_format: "research synthesis"
editorial_timing: {"lane":"regular_hourly","scheduled_at":"2026-08-31T09:00:00+03:00","real_news_delta":"The shared name makes search results and benchmark summaries look comparable even though the evaluated tasks and operational claims are different."}
schema_version: "newruntime-agent-readable-v0.2"
stable_id: "post:two-extractbenches-different-systems"
status: "published"
visuals: []
editorial_provenance: {"schema_version":"newruntime-editorial-copy-v1","content_status":"source_grounded_final","final_copy_sha256":"sha256:7126ef31a7b2ee37bf296782717621a321dfebee843a0382ad5cf080eb2c4194","reviewed_at":"2026-08-27T16:57:21Z","source_evidence_count":2,"verified_claim_count":2,"site_analysis_schema_version":"newruntime-site-analysis-v1","site_object_kind":"field_note","observed_fact_count":2,"implication_count":1,"watch_condition_count":1,"related_record_count":0}
analysis: {"schema_version":"newruntime-site-analysis-v1","object_kind":"field_note","thesis":"The two projects named ExtractBench measure incompatible document-extraction questions and their headline scores must not be compared as one ranking.","observed_facts":[{"text":"LlamaIndex ExtractBench contains 370 enterprise documents and 4,869 pages across eight domains and 67 document types.","source_urls":["https://arxiv.org/abs/2607.29677"]},{"text":"Contextual AI's ExtractBench evaluates complex structured extraction with a different document set, schema regime, and field-level methodology.","source_urls":["https://arxiv.org/abs/2602.12247"]}],"mechanism":"A valid comparison must align datasets, schema breadth, repeated-record cardinality, grounding requirements, scorer definitions, model versions, and inference settings before interpreting a score difference.","why_now":"The shared name makes search results and benchmark summaries look comparable even though the evaluated tasks and operational claims are different.","implications":["Name the benchmark owner and version with every score, preserve the task contract, and compare systems only inside the same harness or through a controlled cross-run."],"evidence_boundary":"Both benchmarks provide useful evidence, but neither result can be translated directly into the other's leaderboard or into a universal claim that document extraction is solved.","watch_conditions":["Revise the comparison only after a common system is rerun on both public harnesses with pinned versions and matched reporting of accuracy, completeness, grounding, cost, and failures."],"related_records":[],"new_branch_reason":"No earlier New Runtime analytical record distinguishes these two same-named benchmark systems."}
routes: {"html":"https://newruntime.com/posts/two-extractbenches-different-systems/","markdown":"https://newruntime.com/posts/two-extractbenches-different-systems.md","json":"https://newruntime.com/posts/two-extractbenches-different-systems.json"}
---

# Two ExtractBenches Measure Different Systems

## Retrieval answer

The two projects named ExtractBench measure incompatible document-extraction questions and their headline scores must not be compared as one ranking. Both benchmarks provide useful evidence, but neither result can be translated directly into the other's leaderboard or into a universal claim that document extraction is solved.

The two projects named ExtractBench measure incompatible document-extraction questions and their headline scores must not be compared as one ranking.

The shared name makes search results and benchmark summaries look comparable even though the evaluated tasks and operational claims are different. LlamaIndex ExtractBench contains 370 enterprise documents and 4,869 pages across eight domains and 67 document types. Contextual AI's ExtractBench evaluates complex structured extraction with a different document set, schema regime, and field-level methodology.

A valid comparison must align datasets, schema breadth, repeated-record cardinality, grounding requirements, scorer definitions, model versions, and inference settings before interpreting a score difference. Name the benchmark owner and version with every score, preserve the task contract, and compare systems only inside the same harness or through a controlled cross-run. This opens a separate benchmark-comparison branch because no existing New Runtime record distinguishes the two projects.

Both benchmarks provide useful evidence, but neither result can be translated directly into the other's leaderboard or into a universal claim that document extraction is solved. Revise the comparison only after a common system is rerun on both public harnesses with pinned versions and matched reporting of accuracy, completeness, grounding, cost, and failures.
