#OpenModels #ScientificAI #PostTraining #AgentHarness
Arcee’s “Teaching an Open Model to Do Science” is not just a model announcement. The useful part is the operating model around post-training: a 21-run program where each run gets a hypothesis, context, plan, versioned configuration, metrics, trace review, and an explicit decision.
The training target is scientific work: tool use, biological reasoning, and auditable research workflows. Arcee describes fixed checkpoints, held-out environments, verifier review, and a promoted run 120 adapter after the Drug Tool score moved from 70.8% to 81.2% while BioReason held at 0.863. The numbers matter less than the discipline around them: one bounded change at a time, then inspect curves, rollouts, traces, and failure modes before keeping the result.
The deployment shape is also more interesting than a leaderboard. The specialist adapter is plugged into a companion AI Scientist application built around an agentic research harness: orchestration, sandboxed Python execution, fixed /plan, /report, and /hypothesize flows, domain skills, session artifacts, and managed serving.
For New Runtime this is the same pattern as editorial automation at a different scale. A stronger model is useful only when the system can say which run changed, which evals held, which traces were reviewed, and why the adapter was promoted. Without that ledger, “the model got better” is not an engineering statement.
