Evidence-linked trend hypothesis

Scientific agents need artifact evidence

Scientific agents are becoming credible when they produce reproducible procedures, artifacts, citations, rewards, and held-out checks rather than persuasive answers alone.

Current thesis

The useful unit of scientific-agent progress is an evidence package that another researcher can inspect, test, rerun, and revise. Confidence: high. Supported by 4 normalized raw signals. This New Runtime record is an evidence-linked retrieval unit. Use its canonical page, machine-readable representations, dates, scope, and public source URLs to verify the claim before reusing it.

Source ledger

Publishable sources attached to this record.

4 public sources
#SourceRolePublic status
1terminal-bench-science.aisourceprimary receiptsource_urls
2alignment.anthropic.comsourcesupporting receiptsource_urls
3thinkingmachines.aisourcesupporting receiptsource_urls
4parallelai.prosourcesupporting receiptsource_urls

What is changing?

Scientific-agent products are moving from impressive answers toward workflows that leave inspectable evidence. The output may be an analysis, simulation, proof, code artifact, dataset, citation set, trained model, or monitored result, but it needs an external check that does not depend on the agent’s prose.

What evidence supports this pattern?

  • Terminal-Bench-Science uses expert-authored terminal workflows and grades concrete artifacts with reproducible, task-specific tests.
  • Anthropic’s automated alignment researchers separate method proposals, training, capability checks, held-out benchmarks, behavioral audits, and trajectory monitoring inside a bounded research loop.
  • Thinking Machines’ Text-to-SQL work makes curated domain data and a semantically aware reward part of the research system.
  • Parallel packages monitored life-science sources, structured outputs, and citations behind workflow-specific APIs, while still making vendor claims that require independent domain validation.

What should teams do next?

Define the evidence object before selecting the agent. Record the procedure, inputs, environment, cost, artifacts, citations, grader, failure modes, and human judgment that can change the result. Preserve enough state for another researcher to rerun or challenge the work instead of accepting a polished final answer as the unit of progress.

Discovery graph / next reads

Continue through New Runtime

Open the graph
  1. 01related materialAgent skills need behavioral evals, not prose reviewSignal used as evidence for this pattern.
  2. 02related materialVerifiability is an AI product featureSignal used as evidence for this pattern.
  3. 03related materialAgent Evals Become a Discipline Separate from Model EvalsSignal used as evidence for this pattern.
  4. 04related materialAnthropic / Ai Resistant Technical Evaluations: Verification BandwidthSignal used as evidence for this pattern.
  5. 05related materialAI Coding Workflow: From Idea to Verifiable WorkField Note connected to this pattern.

These links are also published in this page’s JSON twin and as typed edges in DiscoveryGraph v1.

Who read this page?Machine requests, hidden until opened

Loading the privacy-safe route aggregate…

Open the JSON contract