Field note
Jeremy Berman's public harness reports 96.2% for Claude Opus 5 on 25 ARC-AGI-3 games, versus a 30.2% model-only result cited by the project.
New Runtime reading: The harness supplies a sandbox, filesystem, action broker, durable logs, and runner policy while the model writes task-specific parsers and search code. The result is preliminary and demonstrates a model-plus-runtime system, not a model-only rank.
Evidence boundary: this item uses the listed public sources and keeps vendor, author, or reporter claims attributed. The queued page is an editorial synthesis, not an independent validation of every reported metric.
