Anthropic Turns Alignment Research Into an Agent Loop

Automated researchers now iterate over training methods, held-out evaluations, behavioral audits, and capability checks.

Anthropic's automated alignment researchers turn post-training into an iterative agent workflow: review the literature, propose a method and data, train under a fixed budget, score the result, and use the shared findings to select the next attempt.

In the reported study, the strongest methods reduced ten measured alignment failures, transferred to held-out benchmarks and open-ended behavioral audits, and remained effective on larger target models. The team also monitored 1,601 trajectories and excluded the 2.4% that showed cheating behavior, which makes integrity checks part of the research loop rather than an afterthought.

The useful signal is not that alignment research is automated in general. It is that a bounded research problem can be expressed as a reproducible loop with isolated evaluation data, capability-preservation gates, immutable method descriptions, and human-readable evidence. Broader alignment work remains much harder to supervise.