{"schema_version":"newruntime-agent-readable-v0.2","type":"trend_pattern","stable_id":"pattern:scientific-agents-need-artifact-evidence","slug":"scientific-agents-need-artifact-evidence","title":"Scientific agents need artifact evidence","description":"Scientific agents are becoming credible when they produce reproducible procedures, artifacts, citations, rewards, and held-out checks rather than persuasive answers alone.","retrieval_nugget":"Scientific agents are becoming credible when they produce reproducible procedures, artifacts, citations, rewards, and held-out checks rather than persuasive answers alone. The useful unit of scientific-agent progress is an evidence package that another researcher can inspect, test, rerun, and revise. Confidence is high.","thesis":"The useful unit of scientific-agent progress is an evidence package that another researcher can inspect, test, rerun, and revise.","status":"published","confidence":"high","first_seen":"2026-08-27","last_verified":"2026-09-01","record_date":"2026-09-01","date_kind":"last_verified","supporting_signals":["tg-2707","tg-2588","tg-2409","tg-1185"],"related_posts":["ai-coding-workflow-verifiable-work"],"counter_evidence":["Many exploratory scientific tasks cannot be reduced to deterministic graders without discarding novelty, interpretation, or domain judgment.","A reproducible artifact can still encode a wrong assumption, biased dataset, incomplete source corpus, or reward that measures the wrong outcome."],"revision_trigger":"Revise the thesis if answer-level agents produce reliable scientific gains without reusable procedures, artifacts, citations, or independent validation across domains.","topics":["science-agents","reproducibility","evals","research-infrastructure"],"source_urls":["https://www.terminal-bench-science.ai/announcement","https://alignment.anthropic.com/2026/automated-alignment-researchers/","https://thinkingmachines.ai/news/putting-task-expertise-into-rl/","https://parallelai.pro/solutions/life-sciences"],"routes":{"html":"https://newruntime.com/patterns/scientific-agents-need-artifact-evidence/","markdown":"https://newruntime.com/patterns/scientific-agents-need-artifact-evidence.md","json":"https://newruntime.com/patterns/scientific-agents-need-artifact-evidence.json"},"source_format":"markdown","next_reads":[{"type":"related_material","path":"/signals/agent-skills-need-behavioral-evals/","reason":"Signal used as evidence for this pattern.","url":"https://newruntime.com/signals/agent-skills-need-behavioral-evals/","title":"Agent skills need behavioral evals, not prose review","media_type":"text/html"},{"type":"related_material","path":"/signals/verifiability-is-an-ai-product-feature/","reason":"Signal used as evidence for this pattern.","url":"https://newruntime.com/signals/verifiability-is-an-ai-product-feature/","title":"Verifiability is an AI product feature","media_type":"text/html"},{"type":"related_material","path":"/signals/agent-evals-become-a-discipline-separate-from-model-evals/","reason":"Signal used as evidence for this pattern.","url":"https://newruntime.com/signals/agent-evals-become-a-discipline-separate-from-model-evals/","title":"Agent Evals Become a Discipline Separate from Model Evals","media_type":"text/html"},{"type":"related_material","path":"/signals/anthropic-ai-resistant-technical-evaluations-verification-bandwidth/","reason":"Signal used as evidence for this pattern.","url":"https://newruntime.com/signals/anthropic-ai-resistant-technical-evaluations-verification-bandwidth/","title":"Anthropic / Ai Resistant Technical Evaluations: Verification Bandwidth","media_type":"text/html"},{"type":"related_material","path":"/posts/ai-coding-workflow-verifiable-work/","reason":"Field Note connected to this pattern.","url":"https://newruntime.com/posts/ai-coding-workflow-verifiable-work/","title":"AI Coding Workflow: From Idea to Verifiable Work","media_type":"text/html"}]}
