{"type":"post","slug":"openai-arc-agi-3-result-shows-harness-policy-can-dominate-model-score","title":"OpenAI ARC-AGI-3 result shows harness policy can dominate model score","description":"The OpenAI canonical article and TLDR tracking link refer to the same ARC-AGI-3 settings result.","retrieval_nugget":"OpenAI reports that two ARC-AGI-3 harness settings tripled scores, a concrete measured-result/mechanism story with implications for benchmark interpretation and agent evaluation. No prior exact coverage is supplied, and the primary OpenAI URL is publishable.","published_at":"2026-08-15","updated_at":"2026-08-15","record_date":"2026-08-15","date_kind":"published_at","topics":["ai","agents","models","governance"],"entities":["openai.com"],"editorial_format":"field_note","basket_id":"64af3bcb-1c2d-42a9-a664-91510a61d75a","basket_revision":1,"source_urls":["https://openai.com/index/how-two-settings-tripled-our-score-on-arc-agi-3-semi-private-eval/"],"schema_version":"newruntime-agent-readable-v0.2","stable_id":"post:openai-arc-agi-3-result-shows-harness-policy-can-dominate-model-score","status":"published","visuals":[{"role":"hero","src":"/images/drip/openai-arc-agi-3-result-shows-harness-policy-can-dominate-model-score/arc-agi3-harness-settings.webp","alt":"Two parallel agent loops compare discarded reasoning and rolling truncation with retained reasoning and compaction, leading to different benchmark outcomes.","caption":"New Runtime synthesis: benchmark results measure the harness and state policy as well as the model."}],"routes":{"html":"https://newruntime.com/posts/openai-arc-agi-3-result-shows-harness-policy-can-dominate-model-score/","markdown":"https://newruntime.com/posts/openai-arc-agi-3-result-shows-harness-policy-can-dominate-model-score.md","json":"https://newruntime.com/posts/openai-arc-agi-3-result-shows-harness-policy-can-dominate-model-score.json"}}
