Evidence-linked trend hypothesis
Scientific agents need artifact evidence
Scientific agents are becoming credible when they produce reproducible procedures, artifacts, citations, rewards, and held-out checks rather than persuasive answers alone.
Source ledger
Publishable sources attached to this record.
| # | Source | Role | Public status |
|---|---|---|---|
| 1 | terminal-bench-science.aisource | primary receipt | source_urls |
| 2 | alignment.anthropic.comsource | supporting receipt | source_urls |
| 3 | thinkingmachines.aisource | supporting receipt | source_urls |
| 4 | parallelai.prosource | supporting receipt | source_urls |
What is changing?
Scientific-agent products are moving from impressive answers toward workflows that leave inspectable evidence. The output may be an analysis, simulation, proof, code artifact, dataset, citation set, trained model, or monitored result, but it needs an external check that does not depend on the agent’s prose.
What evidence supports this pattern?
- Terminal-Bench-Science uses expert-authored terminal workflows and grades concrete artifacts with reproducible, task-specific tests.
- Anthropic’s automated alignment researchers separate method proposals, training, capability checks, held-out benchmarks, behavioral audits, and trajectory monitoring inside a bounded research loop.
- Thinking Machines’ Text-to-SQL work makes curated domain data and a semantically aware reward part of the research system.
- Parallel packages monitored life-science sources, structured outputs, and citations behind workflow-specific APIs, while still making vendor claims that require independent domain validation.
What should teams do next?
Define the evidence object before selecting the agent. Record the procedure, inputs, environment, cost, artifacts, citations, grader, failure modes, and human judgment that can change the result. Preserve enough state for another researcher to rerun or challenge the work instead of accepting a polished final answer as the unit of progress.