Field note
Agent evaluation is acquiring diagnostic surfaces around the score. Mercor publishes sample tasks for Terminal-Bench and related benchmarks, DORA's Quick Check connects delivery practices to concrete improvement areas, and AgentBehavior organizes observed behaviors that can appear during an agent trajectory.
The shared move is from “which model ranks first?” to “which system configuration produced this result?” A meaningful record includes the model version, harness, tools, context budget, cost, reviewer interventions, and any unsafe actions along the way. Without that data, a leaderboard cannot explain whether the gain will survive a different repository or policy.
For production selection, eval output should lead to an operating decision: allow, restrict, route, review, or reject a model-plus-harness combination. Open tasks and behavior traces make that decision auditable. They also make regressions easier to localize when a model, tool, or prompt changes independently.