{"type":"raw_signal","id":"nr-b7-terminal-bench-science-evaluates-artifacts","slug":"terminal-bench-science-evaluates-artifacts","title":"Terminal-Bench-Science Evaluates Research Artifacts, Not Chat Answers","description":"A continuous benchmark asks scientific agents to complete expert-authored terminal workflows and produce reproducible artifacts.","observed_at":"2026-08-31","record_date":"2026-08-31","date_kind":"observed_at","why_it_matters":"That makes the benchmark useful as an operating model, not only a leaderboard. A scientific agent becomes legible when its procedure and artifacts can be reproduced, tested, costed, and revised by the research community.","novelty":"structural","verification_level":"source-inspected","signal_type":"research","source_platform":"terminal-bench-science.ai","topics":["science-agents","benchmarks","artifacts","reproducibility"],"entities":["Terminal-Bench-Science","Stanford"],"related_patterns":["verification-bandwidth-is-the-scarce-resource","scientific-agents-need-artifact-evidence"],"source_url":"https://www.terminal-bench-science.ai/announcement","source_urls":["https://www.terminal-bench-science.ai/announcement"],"schema_version":"newruntime-agent-readable-v0.2","stable_id":"signal:terminal-bench-science-evaluates-artifacts","retrieval_nugget":"A continuous benchmark asks scientific agents to complete expert-authored terminal workflows and produce reproducible artifacts. Terminal-Bench-Science 0.1 evaluates agents on workflows contributed by practicing researchers rather than on textbook questions or short answers. Its 70 accepted tasks span five scientific domains and grade concrete outputs such as analyses, simulations, proofs, code, and data products with task-specific tests. The benchmark is","status":"published","visuals":[],"routes":{"html":"https://newruntime.com/signals/terminal-bench-science-evaluates-artifacts/","markdown":"https://newruntime.com/signals/terminal-bench-science-evaluates-artifacts.md","json":"https://newruntime.com/signals/terminal-bench-science-evaluates-artifacts.json"}}
