{"schema_version":"newruntime-agent-readable-v0.2","type":"post","stable_id":"post:openai-arc-agi-settings-harness","slug":"openai-arc-agi-settings-harness","title":"OpenAI's ARC-AGI-3 Jump Was a Harness Result","description":"OpenAI's ARC-AGI-3 write-up shows why agent benchmarks measure the model plus the runtime harness: retained reasoning and compaction changed both score and token use.","retrieval_nugget":"OpenAI's ARC-AGI-3 write-up shows why agent benchmarks measure the model plus the runtime harness: retained reasoning and compaction changed both score and token use. OpenAI's ARC-AGI-3 write-up is useful because it makes a usually hidden fact explicit: agent benchmarks measure the model and the runtime around the model. The public story is not just \"GPT-5.6 Sol scored better.\"","status":"published","published_at":"2026-07-30","updated_at":"2026-07-30","record_date":"2026-07-30","date_kind":"published_at","topics":["evals","agent-harnesses","context-engineering","api-design","models"],"source_urls":["https://x.com/OpenAI/status/2082616641834422740","https://openai.com/index/how-two-settings-tripled-our-arc-agi-3-scores/"],"visuals":[{"id":"openai-arc-agi-settings-harness-nano-banana","kind":"editorial-diagram","role":"hero","src":"https://newruntime.com/images/posts/openai-arc-agi-settings-harness-nano-banana.webp","alt":"Hand-drawn split diagram where a generic harness drops memory while a Responses API-style harness retains reasoning, compacts context, and produces a score jump.","caption":"OpenAI's ARC-AGI-3 result is a reminder that long-running agent evals measure harness memory and context policy, not only model capability.","credit":"New Runtime synthesis from public source inspection","source_url":"https://openai.com/index/how-two-settings-tripled-our-arc-agi-3-scores/","generated_with":"nano-banana-style-imagegen","width":1600,"height":900,"legend":[{"label":"Generic harness","description":"The benchmark runner discarded private reasoning after each action and used rolling truncation as history grew."},{"label":"Retained reasoning","description":"Passing the previous response ID preserves reasoning across turns and tool calls in the Responses API setup."},{"label":"Compaction","description":"Summarizing long context preserved earlier observations better than dropping the oldest messages."},{"label":"Score jump","description":"OpenAI reports 13.3% on the public set with the official harness and 38.3% with retained reasoning plus compaction."}]}],"telegram_message_id":2780,"telegram_url":"https://t.me/qwgai/2780","telegram_message_ids":[2779,2780],"telegram_delivery_mode":"media_then_text","telegram_media_url":"https://t.me/qwgai/2779","routes":{"html":"https://newruntime.com/posts/openai-arc-agi-settings-harness/","markdown":"https://newruntime.com/posts/openai-arc-agi-settings-harness.md","json":"https://newruntime.com/posts/openai-arc-agi-settings-harness.json"},"source_format":"markdown","next_reads":[{"type":"topic","path":"/topics/api-design/","reason":"Explore the api design topic hub.","url":"https://newruntime.com/topics/api-design/","title":"API Design - New Runtime","media_type":"text/html"},{"type":"topic","path":"/topics/context-engineering/","reason":"Explore the context engineering topic hub.","url":"https://newruntime.com/topics/context-engineering/","title":"Context engineering - New Runtime","media_type":"text/html"},{"type":"related_material","path":"/posts/openai-gpt-5-6-efficiency-stack/","reason":"Shares agent harnesses and context engineering.","url":"https://newruntime.com/posts/openai-gpt-5-6-efficiency-stack/","title":"OpenAI Shows Efficiency Is a Full-Stack Agent Problem","media_type":"text/html"},{"type":"related_material","path":"/posts/agentic-sdlc-software-factory-loop/","reason":"Shares agent harnesses and evals.","url":"https://newruntime.com/posts/agentic-sdlc-software-factory-loop/","title":"A Software Factory Connects Agents Through Verified Outcomes","media_type":"text/html"},{"type":"related_material","path":"/posts/chatgpt-agent-loop-efficiency-stack/","reason":"Shares agent harnesses and context engineering.","url":"https://newruntime.com/posts/chatgpt-agent-loop-efficiency-stack/","title":"ChatGPT Cuts Repeated Work Across The Agent Stack","media_type":"text/html"}]}
