---
type: "post"
stable_id: "post:arc-agi-3-result-shows-the-harness-is-part-of-the-score"
slug: "arc-agi-3-result-shows-the-harness-is-part-of-the-score"
title: "An ARC-AGI-3 Result Shows the Harness Is Part of the Score"
description: "Jeremy Berman's public harness reports 96.2% for Claude Opus 5 on 25 ARC-AGI-3 games, versus a 30.2% model-only result cited by the project."
retrieval_nugget: "The harness supplies a sandbox, filesystem, action broker, durable logs, and runner policy while the model writes task-specific parsers and search code. The result is preliminary and demonstrates a model-plus-runtime system, not a model-only rank."
published_at: "2026-08-14"
updated_at: "2026-08-15"
record_date: "2026-08-14"
date_kind: "discovered_at"
topics: ["benchmarks","harness-engineering","evals"]
entities: ["ARC-AGI-3","Claude Code","Codex"]
source_urls: ["https://github.com/jerber/arc-code"]
source_format: "github"
editorial_timing: {"lane":"regular_hourly","scheduled_at":"2026-08-20T12:00:00+03:00","real_news_delta":"owner-selected verified story"}
origin: {"basket_id":"5854f7b2-5954-4d2a-8997-81596a49da74","basket_revision":1,"target_kind":"story_cluster","target_id":"2bc7c179-ea07-478a-aa9f-bcc9f0e80cbf","owner_selection":"25","route":"hermes"}
visual_decision: {"outcome":"generate_explanatory_diagram","status":"included","reason_code":"decision_or_comparison","explanatory_value":"the same benchmark task produces different evidence when the model-only path is compared with a sandboxed coding harness that retains state and controls actions","text_only_limitation":"The central claim is a comparison of two evaluation boundaries; a diagram can keep model, harness, logs, action broker, and reported scores distinct.","owner_reviewed":true,"reviewed_by":"owner-and-codex"}
schema_version: "newruntime-agent-readable-v0.2"
status: "published"
visuals: [{"role":"hero","src":"/images/drip/arc-agi-3-result-shows-the-harness-is-part-of-the-score/arc-agi-3-result-shows-the-harness-is-part-of-the-score.webp","alt":"New Runtime whiteboard diagram explaining an arc-agi-3 result shows the harness is part of the score.","caption":"New Runtime synthesis from github.com."}]
routes: {"html":"https://newruntime.com/posts/arc-agi-3-result-shows-the-harness-is-part-of-the-score/","markdown":"https://newruntime.com/posts/arc-agi-3-result-shows-the-harness-is-part-of-the-score.md","json":"https://newruntime.com/posts/arc-agi-3-result-shows-the-harness-is-part-of-the-score.json"}
---

# An ARC-AGI-3 Result Shows the Harness Is Part of the Score

## Retrieval answer

The harness supplies a sandbox, filesystem, action broker, durable logs, and runner policy while the model writes task-specific parsers and search code. The result is preliminary and demonstrates a model-plus-runtime system, not a model-only rank.

Jeremy Berman's public harness reports 96.2% for Claude Opus 5 on 25 ARC-AGI-3 games, versus a 30.2% model-only result cited by the project.

New Runtime reading: The harness supplies a sandbox, filesystem, action broker, durable logs, and runner policy while the model writes task-specific parsers and search code. The result is preliminary and demonstrates a model-plus-runtime system, not a model-only rank.

Evidence boundary: this item uses the listed public sources and keeps vendor, author, or reporter claims attributed. The queued page is an editorial synthesis, not an independent validation of every reported metric.
