{"schema_version":"newruntime-agent-readable-v0.2","type":"post","stable_id":"post:proof-leaves-the-executor","slug":"proof-leaves-the-executor","title":"When Proof Leaves the Executor","description":"Agent autonomy becomes operational when completion claims stop being self-authenticating and tests, evaluators, review gates, and writeback authority remain outside the executor.","retrieval_nugget":"Agent autonomy becomes operational when completion claims stop being self-authenticating and tests, evaluators, review gates, and writeback authority remain outside the executor. An agent says it has finished. That sentence is becoming less important. As AI systems receive longer runtimes, broader tools, and permission to revise their own work, a completion message can no longer serve as proof.","status":"published","published_at":"2026-09-08","updated_at":"2026-09-08","updated_at_kind":"published_at_fallback","record_date":"2026-09-08","date_kind":"published_at","topics":["agent-operations","verification","evals","governance","organizational-design"],"source_urls":["https://www.databricks.com/blog/benchmarking-coding-agents-databricks-multi-million-line-codebase","https://www.terminal-bench-science.ai/announcement","https://alignment.anthropic.com/2026/automated-alignment-researchers/","https://thinkingmachines.ai/news/putting-task-expertise-into-rl/","https://www.uber.com/us/en/blog/ureview/","https://cline.bot/blog/recursive-self-improvement-for-coding-agents"],"visuals":[{"id":"proof-leaves-the-executor","kind":"editorial-diagram","role":"hero","src":"https://newruntime.com/images/posts/proof-leaves-the-executor/proof-leaves-the-executor.webp","alt":"Hand-drawn systems diagram showing an AI executor producing attempt documents inside a bounded purple workspace. Independent tests, an evaluator, and accountable human review sit outside the workspace and feed a green acceptance gate. Only an accepted evidence package crosses into the blue system of record.","caption":"New Runtime synthesis: autonomy scales when attempts and acceptance are owned by different layers.","credit":"New Runtime synthesis from the cited public evidence set","source_url":"https://www.databricks.com/blog/benchmarking-coding-agents-databricks-multi-million-line-codebase","generated_with":"gemini-3.1-flash-image","width":1600,"height":900,"legend":[{"label":"Bounded execution","description":"The agent can create, inspect, and revise attempts inside a constrained workspace."},{"label":"Independent evaluation","description":"Held-out tests, task-specific graders, monitors, and human judgment do not inherit the executor's completion claim."},{"label":"External acceptance","description":"A separate gate decides whether the evidence is sufficient for the result to change shared state."},{"label":"Durable record","description":"Only the accepted artifact, evidence, and decision enter the system of record."}]}],"routes":{"html":"https://newruntime.com/posts/proof-leaves-the-executor/","markdown":"https://newruntime.com/posts/proof-leaves-the-executor.md","json":"https://newruntime.com/posts/proof-leaves-the-executor.json"},"source_format":"markdown","next_reads":[{"type":"topic","path":"/topics/evals/","reason":"Explore the evals topic hub.","url":"https://newruntime.com/topics/evals/","title":"Agent evals - New Runtime","media_type":"text/html"},{"type":"topic","path":"/topics/governance/","reason":"Explore the governance topic hub.","url":"https://newruntime.com/topics/governance/","title":"Governance - New Runtime","media_type":"text/html"},{"type":"related_material","path":"/posts/langchain-reviewbench-review-agent-evals/","reason":"Shares evals and verification.","url":"https://newruntime.com/posts/langchain-reviewbench-review-agent-evals/","title":"ReviewBench Turns Code Review Into An Agent Eval","media_type":"text/html"},{"type":"related_material","path":"/posts/ramp-agentic-risk-operations/","reason":"Shares evals and governance.","url":"https://newruntime.com/posts/ramp-agentic-risk-operations/","title":"Ramp Separates Agent Reasoning From Risk Decisions","media_type":"text/html"},{"type":"related_material","path":"/posts/anthropic-claude-cryptographic-weaknesses/","reason":"Shares evals and verification.","url":"https://newruntime.com/posts/anthropic-claude-cryptographic-weaknesses/","title":"Claude Mythos Moves Cryptanalysis Into the Verification Bottleneck","media_type":"text/html"}]}
