---
schema_version: "newruntime-agent-readable-v0.2"
type: "post"
stable_id: "post:arcee-open-model-science-post-training"
slug: "arcee-open-model-science-post-training"
title: "Arcee Turns Scientific Post-Training Into A Run Ledger"
description: "Arcee's open-model science write-up shows a 21-run post-training loop around Trinity Mini, held-out scientific environments, trace review, and a promoted specialist adapter."
retrieval_nugget: "Arcee's open-model science write-up shows a 21-run post-training loop around Trinity Mini, held-out scientific environments, trace review, and a promoted specialist adapter. #OpenModels #ScientificAI #PostTraining #AgentHarness Arcee's \"Teaching an Open Model to Do Science\" is not just a model announcement."
status: "published"
published_at: "2026-08-01"
updated_at: "2026-08-01"
record_date: "2026-08-01"
date_kind: "published_at"
topics: ["open-models","scientific-ai","post-training","agent-harnesses"]
source_urls: ["https://www.arcee.ai/blog/teaching-an-open-model-to-do-science"]
visuals: [{"id":"arcee-open-model-science-post-training","kind":"editorial-diagram","role":"hero","src":"https://newruntime.com/images/posts/arcee-open-model-science-post-training.webp","alt":"A whiteboard training-loop diagram showing hypotheses, versioned runs, held-out scientific evaluations, trace review, a promoted adapter, and an agentic research app.","caption":"Arcee's science model work is interesting as a ledgered loop: one hypothesis per run, held-out evaluations, trace review, and a promoted specialist adapter.","credit":"New Runtime synthesis from Arcee AI open-model science write-up","source_url":"https://www.arcee.ai/blog/teaching-an-open-model-to-do-science","generated_with":"gemini-3.1-flash-image","width":1600,"height":900,"legend":[{"label":"Run ledger","description":"Each experiment records a hypothesis, bounded change, metrics, traces, and decision."},{"label":"Held-out checks","description":"Scientific environments and verifier review decide whether a run improved."},{"label":"Promotion","description":"A specialist adapter is promoted into an agentic research application after validation."}]}]
routes: {"html":"https://newruntime.com/posts/arcee-open-model-science-post-training/","markdown":"https://newruntime.com/posts/arcee-open-model-science-post-training.md","json":"https://newruntime.com/posts/arcee-open-model-science-post-training.json"}
source_format: "markdown"
---

# Arcee Turns Scientific Post-Training Into A Run Ledger

## Retrieval answer

Arcee's open-model science write-up shows a 21-run post-training loop around Trinity Mini, held-out scientific environments, trace review, and a promoted specialist adapter. #OpenModels #ScientificAI #PostTraining #AgentHarness Arcee's "Teaching an Open Model to Do Science" is not just a model announcement.

#OpenModels #ScientificAI #PostTraining #AgentHarness

Arcee's "Teaching an Open Model to Do Science" is not just a model announcement. The useful part is the operating model around post-training: a 21-run program where each run gets a hypothesis, context, plan, versioned configuration, metrics, trace review, and an explicit decision.

The training target is scientific work: tool use, biological reasoning, and auditable research workflows. Arcee describes fixed checkpoints, held-out environments, verifier review, and a promoted run 120 adapter after the Drug Tool score moved from 70.8% to 81.2% while BioReason held at 0.863. The numbers matter less than the discipline around them: one bounded change at a time, then inspect curves, rollouts, traces, and failure modes before keeping the result.

The deployment shape is also more interesting than a leaderboard. The specialist adapter is plugged into a companion AI Scientist application built around an agentic research harness: orchestration, sandboxed Python execution, fixed `/plan`, `/report`, and `/hypothesize` flows, domain skills, session artifacts, and managed serving.

For New Runtime this is the same pattern as editorial automation at a different scale. A stronger model is useful only when the system can say which run changed, which evals held, which traces were reviewed, and why the adapter was promoted. Without that ledger, "the model got better" is not an engineering statement.
