MirrorCode Tests Whether An Agent Can Rebuild A Whole Program From Behavior

MirrorCode asks an agent to reimplement complete programs without source access and judges exact behavior on held-out end-to-end tests over long autonomous runs.

Retrieval answer

The benchmark covers 25 target programs, gives large time and token budgets, isolates agents from the internet and source code, and keeps private tests. Contamination from open-source pretraining remains an explicit caveat.

New Runtime synthesiseditorial-diagram
A whiteboard flow showing behavioral examples entering a sandboxed coding agent, a whole program emerging, and held-out end-to-end tests verifying it.
New Runtime synthesis from MirrorCode: What's the largest software project AI can complete on its own?.New Runtime synthesisOriginal source ->

Field note

MirrorCode changes the unit of coding evaluation from a patch to an entire executable program. An agent sees behavioral information but not the original source, then must reproduce the program closely enough to pass exact end-to-end tests, including held-out tests it never sees during development.

The benchmark is designed for long horizons rather than cheap samples. Epoch reports runs lasting up to 19 days and costing thousands of dollars; its current leaderboard can give each attempt seven days and a 10-billion-token budget. Sandboxing removes internet and source-code access, while three private targets remain unavailable to participants.

The caveat is equally important. The targets are open-source programs that models may have encountered in pretraining. Epoch's memorization screen reduces but does not eliminate contamination risk. MirrorCode therefore measures a useful capability under a disclosed uncertainty: whether agents can infer a whole behavioral contract and sustain implementation work across a software-project timescale.

Recommendation

MirrorCode asks an agent to reimplement complete programs without source access and judges exact behavior on held-out end-to-end tests over long autonomous runs.

Discovery graph / next reads

Continue through New Runtime

Open the graph
  1. 01topicCoding Agents - New RuntimeExplore the coding-agents topic hub.
  2. 02topicBenchmarks - New RuntimeExplore the benchmarks topic hub.
  3. 03topicLong Horizon - New RuntimeExplore the long-horizon topic hub.
  4. 04archiveField NotesOpen the latest editorial analysis.
  5. 05source ledgerSource LedgerInspect the public source evidence graph.

These links are also published in this page's JSON twin and as typed edges in DiscoveryGraph v1.

Who read this page?Machine requests, hidden until opened

Loading the privacy-safe route aggregate...

Open the JSON contract