Agent Evaluation Is Becoming a Control Stack

New engineering and research material connects repeatable tasks, trajectory review, public eval validation, and misalignment checks.

Agent evaluation is becoming a control stack that surrounds execution rather than a score attached after the fact.

OpenAI's repetitive-work automation, Airbnb's eval-driven development, public-eval validation, and research on accidental chain-of-thought grading point to different layers of the same problem. Teams need representative tasks, observable trajectories, graders, environment boundaries, and a way to challenge the grader itself.

The operational consequence is to treat eval artifacts as production infrastructure. Preserve the task version, environment, model, tools, evidence, and failure category. A benchmark gain is trustworthy only within that contract, and a passing score should not grant an agent broader authority than the evaluated environment.