---
type: "post"
slug: "hours-visible-autonomous-weeks-not"
title: "Hours Are Visible; Autonomous Weeks Are Not"
description: "Current evidence supports longer bounded agent runs, but not reliable unattended work across weeks."
retrieval_nugget: "Published systems now make hour-scale agent work visible, but they do not yet establish reliable unattended autonomy across days or weeks. The cited systems demonstrate bounded long trajectories and research programs; they do not demonstrate robust autonomous execution across weeks in an open changing environment."
published_at: "2026-08-27"
updated_at: "2026-08-27"
record_date: "2026-08-27"
date_kind: "published_at"
topics: ["new-feature"]
entities: []
source_urls: ["https://arxiv.org/abs/2608.23283","https://arxiv.org/abs/2608.19799","https://deepmind.google/discover/blog/eve-towards-long-horizon-agents/"]
source_format: "research synthesis"
editorial_timing: {"lane":"regular_hourly","scheduled_at":"2026-08-31T08:00:00+03:00","real_news_delta":"Agent products increasingly advertise duration, while the evidence still mixes elapsed time, tool volume, supervision, recovery, and genuine autonomy."}
schema_version: "newruntime-agent-readable-v0.2"
stable_id: "post:hours-visible-autonomous-weeks-not"
status: "published"
visuals: []
editorial_provenance: {"schema_version":"newruntime-editorial-copy-v1","content_status":"source_grounded_final","final_copy_sha256":"sha256:cd9278823db785b7fa438298da0f946c4bc0365c6e2d5ea9f252eef71ecfb16a","reviewed_at":"2026-08-27T16:57:21Z","source_evidence_count":1,"verified_claim_count":2,"site_analysis_schema_version":"newruntime-site-analysis-v1","site_object_kind":"field_note","observed_fact_count":2,"implication_count":1,"watch_condition_count":1,"related_record_count":1}
analysis: {"schema_version":"newruntime-site-analysis-v1","object_kind":"field_note","thesis":"Published systems now make hour-scale agent work visible, but they do not yet establish reliable unattended autonomy across days or weeks.","observed_facts":[{"text":"Apodex reports an endurance example lasting 82.4 minutes with 864 recorded steps and 462 tool calls.","source_urls":["https://arxiv.org/abs/2608.23283"]},{"text":"Its current coordination state is run-scoped and is not atomically checkpointed with the workspace for crash-consistent restart or historical rollback.","source_urls":["https://arxiv.org/abs/2608.23283"]}],"mechanism":"Long work survives by externalizing authoritative state, dependencies, artifacts, verification, checkpoints, and recovery instead of relying on an ever-growing conversational transcript.","why_now":"Agent products increasingly advertise duration, while the evidence still mixes elapsed time, tool volume, supervision, recovery, and genuine autonomy.","implications":["Report duration and autonomy separately, and require durable task state plus restart tests before calling a workflow week-grade."],"evidence_boundary":"The cited systems demonstrate bounded long trajectories and research programs; they do not demonstrate robust autonomous execution across weeks in an open changing environment.","watch_conditions":["Change this conclusion when a public system survives process and machine failure, resumes from a consistent checkpoint, and completes multi-day work under independent verification."],"related_records":[{"url":"https://newruntime.com/patterns/goal-scoped-loops-replace-manual-continuation","relation":"The duration claim becomes useful only when completion and recovery remain external contracts."}]}
routes: {"html":"https://newruntime.com/posts/hours-visible-autonomous-weeks-not/","markdown":"https://newruntime.com/posts/hours-visible-autonomous-weeks-not.md","json":"https://newruntime.com/posts/hours-visible-autonomous-weeks-not.json"}
---

# Hours Are Visible; Autonomous Weeks Are Not

## Retrieval answer

Published systems now make hour-scale agent work visible, but they do not yet establish reliable unattended autonomy across days or weeks. The cited systems demonstrate bounded long trajectories and research programs; they do not demonstrate robust autonomous execution across weeks in an open changing environment.

Published systems now make hour-scale agent work visible, but they do not yet establish reliable unattended autonomy across days or weeks.

Agent products increasingly advertise duration, while the evidence still mixes elapsed time, tool volume, supervision, recovery, and genuine autonomy. Apodex reports an endurance example lasting 82.4 minutes with 864 recorded steps and 462 tool calls. Its current coordination state is run-scoped and is not atomically checkpointed with the workspace for crash-consistent restart or historical rollback.

Long work survives by externalizing authoritative state, dependencies, artifacts, verification, checkpoints, and recovery instead of relying on an ever-growing conversational transcript. Report duration and autonomy separately, and require durable task state plus restart tests before calling a workflow week-grade. This extends the [related New Runtime pattern](https://newruntime.com/patterns/goal-scoped-loops-replace-manual-continuation/): The duration claim becomes useful only when completion and recovery remain external contracts.

The cited systems demonstrate bounded long trajectories and research programs; they do not demonstrate robust autonomous execution across weeks in an open changing environment. Change this conclusion when a public system survives process and machine failure, resumes from a consistent checkpoint, and completes multi-day work under independent verification.
