Field note
The OpenAI canonical article and TLDR tracking link refer to the same ARC-AGI-3 settings result.
Why it matters
OpenAI reports that two ARC-AGI-3 harness settings tripled scores, a concrete measured-result/mechanism story with implications for benchmark interpretation and agent evaluation. No prior exact coverage is supplied, and the primary OpenAI URL is publishable.
New Runtime view
ARC-AGI-3 shows benchmark results measure the harness as well as the model. State policy is part of capability.
Mechanism: Retained reasoning plus context summarization/compaction changes what state survives between attempts.
Architectural boundary: Base model capability is separated from harness memory and runtime policy.
Measured consequence: OpenAI reports the two settings tripled the score on ARC-AGI-3 semi-private eval.
What remains open
- Semi-private eval limits independent reproduction.
- Harness settings can be task-format specific.
- Model comparisons without runtime settings are incomplete.
Sources
- <https://openai.com/index/how-two-settings-tripled-our-score-on-arc-agi-3-semi-private-eval/>
