OpenAI ARC-AGI-3 result shows harness policy can dominate model score

The OpenAI canonical article and TLDR tracking link refer to the same ARC-AGI-3 settings result.

Retrieval answer

OpenAI reports that two ARC-AGI-3 harness settings tripled scores, a concrete measured-result/mechanism story with implications for benchmark interpretation and agent evaluation. No prior exact coverage is supplied, and the primary OpenAI URL is publishable.

New Runtime synthesiseditorial-diagram
Two parallel agent loops compare discarded reasoning and rolling truncation with retained reasoning and compaction, leading to different benchmark outcomes.
New Runtime synthesis: benchmark results measure the harness and state policy as well as the model.New Runtime synthesisOriginal source ->

Field note

The OpenAI canonical article and TLDR tracking link refer to the same ARC-AGI-3 settings result.

Why it matters

OpenAI reports that two ARC-AGI-3 harness settings tripled scores, a concrete measured-result/mechanism story with implications for benchmark interpretation and agent evaluation. No prior exact coverage is supplied, and the primary OpenAI URL is publishable.

New Runtime view

ARC-AGI-3 shows benchmark results measure the harness as well as the model. State policy is part of capability.

Mechanism: Retained reasoning plus context summarization/compaction changes what state survives between attempts.

Architectural boundary: Base model capability is separated from harness memory and runtime policy.

Measured consequence: OpenAI reports the two settings tripled the score on ARC-AGI-3 semi-private eval.

What remains open

  • Semi-private eval limits independent reproduction.
  • Harness settings can be task-format specific.
  • Model comparisons without runtime settings are incomplete.

Sources

  • <https://openai.com/index/how-two-settings-tripled-our-score-on-arc-agi-3-semi-private-eval/>

Recommendation

The OpenAI canonical article and TLDR tracking link refer to the same ARC-AGI-3 settings result.

Discovery graph / next reads

Continue through New Runtime

Open the graph
  1. 01topicAi - New RuntimeExplore the ai topic hub.
  2. 02topicAgents - New RuntimeExplore the agents topic hub.
  3. 03topicModels - New RuntimeExplore the models topic hub.
  4. 04archiveField NotesOpen the latest editorial analysis.
  5. 05source ledgerSource LedgerInspect the public source evidence graph.

These links are also published in this page's JSON twin and as typed edges in DiscoveryGraph v1.

Who read this page?Machine requests, hidden until opened

Loading the privacy-safe route aggregate...

Open the JSON contract