---
type: "raw_signal"
id: "nr-b7-terminal-bench-science-evaluates-artifacts"
slug: "terminal-bench-science-evaluates-artifacts"
title: "Terminal-Bench-Science Evaluates Research Artifacts, Not Chat Answers"
description: "A continuous benchmark asks scientific agents to complete expert-authored terminal workflows and produce reproducible artifacts."
observed_at: "2026-08-31"
record_date: "2026-08-31"
date_kind: "observed_at"
why_it_matters: "That makes the benchmark useful as an operating model, not only a leaderboard. A scientific agent becomes legible when its procedure and artifacts can be reproduced, tested, costed, and revised by the research community."
novelty: "structural"
verification_level: "source-inspected"
signal_type: "research"
source_platform: "terminal-bench-science.ai"
topics: ["science-agents","benchmarks","artifacts","reproducibility"]
entities: ["Terminal-Bench-Science","Stanford"]
related_patterns: ["verification-bandwidth-is-the-scarce-resource","scientific-agents-need-artifact-evidence"]
source_url: "https://www.terminal-bench-science.ai/announcement"
source_urls: ["https://www.terminal-bench-science.ai/announcement"]
schema_version: "newruntime-agent-readable-v0.2"
stable_id: "signal:terminal-bench-science-evaluates-artifacts"
retrieval_nugget: "A continuous benchmark asks scientific agents to complete expert-authored terminal workflows and produce reproducible artifacts. Terminal-Bench-Science 0.1 evaluates agents on workflows contributed by practicing researchers rather than on textbook questions or short answers. Its 70 accepted tasks span five scientific domains and grade concrete outputs such as analyses, simulations, proofs, code, and data products with task-specific tests. The benchmark is"
status: "published"
visuals: []
routes: {"html":"https://newruntime.com/signals/terminal-bench-science-evaluates-artifacts/","markdown":"https://newruntime.com/signals/terminal-bench-science-evaluates-artifacts.md","json":"https://newruntime.com/signals/terminal-bench-science-evaluates-artifacts.json"}
---

# Terminal-Bench-Science Evaluates Research Artifacts, Not Chat Answers

## Retrieval answer

A continuous benchmark asks scientific agents to complete expert-authored terminal workflows and produce reproducible artifacts. Terminal-Bench-Science 0.1 evaluates agents on workflows contributed by practicing researchers rather than on textbook questions or short answers. Its 70 accepted tasks span five scientific domains and grade concrete outputs such as analyses, simulations, proofs, code, and data products with task-specific tests. The benchmark is

Terminal-Bench-Science 0.1 evaluates agents on workflows contributed by practicing researchers rather than on textbook questions or short answers. Its 70 accepted tasks span five scientific domains and grade concrete outputs such as analyses, simulations, proofs, code, and data products with task-specific tests.

The benchmark is designed as a continuous system: researchers propose workflows, domain and technical reviewers inspect them, difficult tasks enter versioned releases, and results can be regraded or rerun as agents change. The launch reports a 30% resolution rate for the strongest evaluated system, leaving substantial room for improvement.

That makes the benchmark useful as an operating model, not only a leaderboard. A scientific agent becomes legible when its procedure and artifacts can be reproduced, tested, costed, and revised by the research community.
