---
type: "post"
slug: "openai-arc-agi-3-result-shows-harness-policy-can-dominate-model-score"
title: "OpenAI ARC-AGI-3 result shows harness policy can dominate model score"
description: "The OpenAI canonical article and TLDR tracking link refer to the same ARC-AGI-3 settings result."
retrieval_nugget: "OpenAI reports that two ARC-AGI-3 harness settings tripled scores, a concrete measured-result/mechanism story with implications for benchmark interpretation and agent evaluation. No prior exact coverage is supplied, and the primary OpenAI URL is publishable."
published_at: "2026-08-15"
updated_at: "2026-08-15"
record_date: "2026-08-15"
date_kind: "published_at"
topics: ["ai","agents","models","governance"]
entities: ["openai.com"]
editorial_format: "field_note"
basket_id: "64af3bcb-1c2d-42a9-a664-91510a61d75a"
basket_revision: 1
source_urls: ["https://openai.com/index/how-two-settings-tripled-our-score-on-arc-agi-3-semi-private-eval/"]
schema_version: "newruntime-agent-readable-v0.2"
stable_id: "post:openai-arc-agi-3-result-shows-harness-policy-can-dominate-model-score"
status: "published"
visuals: [{"role":"hero","src":"/images/drip/openai-arc-agi-3-result-shows-harness-policy-can-dominate-model-score/arc-agi3-harness-settings.webp","alt":"Two parallel agent loops compare discarded reasoning and rolling truncation with retained reasoning and compaction, leading to different benchmark outcomes.","caption":"New Runtime synthesis: benchmark results measure the harness and state policy as well as the model."}]
routes: {"html":"https://newruntime.com/posts/openai-arc-agi-3-result-shows-harness-policy-can-dominate-model-score/","markdown":"https://newruntime.com/posts/openai-arc-agi-3-result-shows-harness-policy-can-dominate-model-score.md","json":"https://newruntime.com/posts/openai-arc-agi-3-result-shows-harness-policy-can-dominate-model-score.json"}
---

# OpenAI ARC-AGI-3 result shows harness policy can dominate model score

## Retrieval answer

OpenAI reports that two ARC-AGI-3 harness settings tripled scores, a concrete measured-result/mechanism story with implications for benchmark interpretation and agent evaluation. No prior exact coverage is supplied, and the primary OpenAI URL is publishable.

The OpenAI canonical article and TLDR tracking link refer to the same ARC-AGI-3 settings result.

## Why it matters

OpenAI reports that two ARC-AGI-3 harness settings tripled scores, a concrete measured-result/mechanism story with implications for benchmark interpretation and agent evaluation. No prior exact coverage is supplied, and the primary OpenAI URL is publishable.

## New Runtime view

ARC-AGI-3 shows benchmark results measure the harness as well as the model. State policy is part of capability.

Mechanism: Retained reasoning plus context summarization/compaction changes what state survives between attempts.

Architectural boundary: Base model capability is separated from harness memory and runtime policy.

Measured consequence: OpenAI reports the two settings tripled the score on ARC-AGI-3 semi-private eval.

## What remains open

- Semi-private eval limits independent reproduction.
- Harness settings can be task-format specific.
- Model comparisons without runtime settings are incomplete.

## Sources

- <https://openai.com/index/how-two-settings-tripled-our-score-on-arc-agi-3-semi-private-eval/>
