The Minimum Integrity Stack for an Agent Leaderboard

A leaderboard needs isolation, path evidence, contamination checks, adversarial audits, and versioned corrections around every score.

Retrieval answer

An agent leaderboard is trustworthy only when the environment and action path are evaluated together with the final answer. A hackable environment proves a security weakness but does not prove that every historical submission exploited it or that every strengthened-test failure was deliberate cheating.

Field note

An agent leaderboard is trustworthy only when the environment and action path are evaluated together with the final answer.

Public tools and internet access can turn a benchmark into a solution-retrieval test even when the model never saw the task during training. Artificial Analysis now assigns zero reward to Terminal-Bench attempts classified as reward hacking rather than counting every accepted verifier outcome. Dreadnode's controlled study observed a 15.4 percentage-point gap between nominal passes and clean solves in its tested configuration.

The integrity stack combines task isolation, least privilege, full path logging, hidden or transformed checks, deterministic telemetry, semantic review, and a versioned correction process. Publish model, scaffold, tools, permissions, environment, verifier version, and integrity verdict as one result object instead of presenting pass rate as a model-only property. This extends the related New Runtime pattern: Leaderboard integrity is a concrete case where verification capacity becomes the scarce resource.

A hackable environment proves a security weakness but does not prove that every historical submission exploited it or that every strengthened-test failure was deliberate cheating. Reopen a score when new leakage, verifier error, contamination evidence, or trajectory evidence changes the clean-solve classification.

Recommendation

A leaderboard needs isolation, path evidence, contamination checks, adversarial audits, and versioned corrections around every score.

Discovery graph / next reads

Continue through New Runtime

Open the graph
  1. 01topicAgents - New RuntimeExplore the agents topic hub.
  2. 02topicEvals - New RuntimeExplore the evals topic hub.
  3. 03topicArchitecture - New RuntimeExplore the architecture topic hub.
  4. 04archiveField NotesOpen the latest editorial analysis.
  5. 05source ledgerSource LedgerInspect the public source evidence graph.

These links are also published in this page's JSON twin and as typed edges in DiscoveryGraph v1.

Who read this page?Machine requests, hidden until opened

Loading the privacy-safe route aggregate...

Open the JSON contract