Field note
METR's note matters because misalignment incidents are becoming evidence-management problems, not only PR or system-card anecdotes.
The useful frame is independent propensity investigation after an incident. A single bad episode does not automatically prove a stable model tendency, but it can define a testable question. What evidence should be preserved? Which prompts, tools, logs, model versions, and environmental details are needed? Who can inspect them without turning the incident into a vendor-controlled story?
For New Runtime, this connects directly to agent operations. As agents gain more tools and longer runtime, incident review has to look more like forensics: chain of custody, replayable conditions, explicit scope, and outside review. Without that, every concerning behavior becomes either under-interpreted as a one-off or over-interpreted as a universal model trait.
